arXiv:2609.02092v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet…
Tag: AI
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual…
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key–value (KV) cache during decoding, which consumes substantial…
MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
arXiv:2609.02060v1 Announce Type: new Abstract: Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence,…
Anthropic’s Claude Fable 5.1 promises better coding and research at up to 45 percent less
Anthropic launches Claude Fable 5.1 and Mythos 5.1, its most capable AI models yet. Fable 5.1 doubles its predecessor’s score on Terminal-Bench-Science…
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
arXiv:2609.02057v1 Announce Type: new Abstract: Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits…
Anthropic’s new Fable release is cheaper, less restrictive
Fable 5.1 includes changes meant to reduce token cost and false-positive restrictions from the model’s safeguards.
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from…
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
arXiv:2609.01924v1 Announce Type: new Abstract: Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard…
When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor
arXiv:2609.01985v1 Announce Type: new Abstract: As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization…
