arXiv:2609.09928v1 Announce Type: new Abstract: Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit…
Grounded Evaluation and Repair for NL-to-PDDL Problem Generation
arXiv:2609.09898v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning…
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
arXiv:2609.09875v1 Announce Type: new Abstract: Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion…
Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications
arXiv:2609.09885v1 Announce Type: new Abstract: This paper studies unmanned aerial vehicle (UAV)-mouted reconfigurable intelligent surface (RIS)-assisted…
Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its…
Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models
arXiv:2609.09925v1 Announce Type: new Abstract: Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of…
Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and…
Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format
arXiv:2609.09882v1 Announce Type: new Abstract: Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these…
AI News Brief Hourly Summary 2026-09-11 09h : 12 posts
12 posts published in the last hour 06:33UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model 06:33Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward 06:33Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks 06:33The…
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
arXiv:2609.09815v1 Announce Type: new Abstract: Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting…
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
arXiv:2609.09776v1 Announce Type: new Abstract: Frontier gains in language-model reasoning come from reinforcement learning on reasoning traces and are…
Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks
arXiv:2609.09774v1 Announce Type: new Abstract: Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine…
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents
arXiv:2609.09853v1 Announce Type: new Abstract: LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no…
Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields
arXiv:2609.09864v1 Announce Type: new Abstract: Affective computing has largely followed an individual-state paradigm, extracting discrete emotion labels…
Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation
arXiv:2609.09702v1 Announce Type: new Abstract: Candidate decision correctness and rationale grounding are different objectives. We examine…
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though…
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
arXiv:2609.09678v1 Announce Type: new Abstract: Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or…
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
arXiv:2609.09754v1 Announce Type: new Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic…
