arXiv:2608.23651v1 Announce Type: cross Abstract: Agent harnesses record a failed tool call and its error message in the transcript and ask the model to…
Tag: cs.AI updates on arXiv.org
ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents
arXiv:2608.23635v1 Announce Type: cross Abstract: Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to…
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
arXiv:2608.23663v2 Announce Type: cross Abstract: Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device…
When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs
arXiv:2608.23623v1 Announce Type: cross Abstract: Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify…
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
arXiv:2608.23660v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural…
Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap
arXiv:2608.23658v1 Announce Type: cross Abstract: An LLM serving engine sizes its key-value (KV) cache once, at startup, permanently setting aside a…
Macro-Operator Generation and Predicate Selection for TAMP Operator Learning
arXiv:2608.23629v1 Announce Type: cross Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning…
REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring
arXiv:2608.23611v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer new opportunities for automated code refactoring. However, generated…
Identifying Latent Declarative Representations of Code for Assisting Repository Migration
arXiv:2608.23619v1 Announce Type: cross Abstract: Legacy software repositories embed decades of domain knowledge in undocumented code, making…
SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
arXiv:2608.24870v1 Announce Type: new Abstract: Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for…
