arXiv:2609.02074v1 Announce Type: new Abstract: Planning is a central capability that enables agents to decompose complex long-horizon tasks into…
Category: cs.AI updates on arXiv.org
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
arXiv:2609.02067v1 Announce Type: new Abstract: Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another’s…
READY or Not: Reliable Enterprise Agent Deployment
arXiv:2609.02095v1 Announce Type: new Abstract: An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent…
Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems
arXiv:2609.02092v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet…
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual…
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key–value (KV) cache during decoding, which consumes substantial…
MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
arXiv:2609.02060v1 Announce Type: new Abstract: Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence,…
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
arXiv:2609.02057v1 Announce Type: new Abstract: Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits…
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from…
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
arXiv:2609.01924v1 Announce Type: new Abstract: Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard…
