XPeng Inc. has commissioned production lines for humanoid robots at its manufacturing facility and completed the line production of its first advanced…
Author: script
Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades
arXiv:2609.01345v2 Announce Type: replace Abstract: Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to…
AI News Brief Hourly Summary 2026-09-08 05h : 12 posts
12 posts published in the last hour 02:32Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors 02:32FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents 02:32AI Revealed Preferences 02:32$A^2E$ : An End-to-End Agent Auditing Engine 02:32NxN E-valuation:…
Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
arXiv:2608.23873v3 Announce Type: replace Abstract: Everything a language model sees is tokens. The serving stack knows what each span is — user input,…
FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
arXiv:2608.29372v2 Announce Type: replace Abstract: Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out…
AI Revealed Preferences
arXiv:2608.26178v2 Announce Type: replace Abstract: There is growing interest in whether language models have stable preferences, for technical, safety,…
$A^2E$ : An End-to-End Agent Auditing Engine
arXiv:2608.07346v3 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential…
NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
arXiv:2608.06621v3 Announce Type: replace Abstract: We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a…
KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation
arXiv:2607.27231v2 Announce Type: replace Abstract: Modern AI systems depend on specialized accelerator kernels, whose development is complicated by…
EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
arXiv:2607.09773v2 Announce Type: replace Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially…
