arXiv:2608.23635v1 Announce Type: cross Abstract: Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to…
Tag: cs.AI updates on arXiv.org
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
arXiv:2608.23660v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural…
Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap
arXiv:2608.23658v1 Announce Type: cross Abstract: An LLM serving engine sizes its key-value (KV) cache once, at startup, permanently setting aside a…
REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring
arXiv:2608.23611v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer new opportunities for automated code refactoring. However, generated…
When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs
arXiv:2608.23623v1 Announce Type: cross Abstract: Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify…
Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal
arXiv:2608.23616v1 Announce Type: cross Abstract: An AI agent’s rebuild is only as good as the process that produced it. Prior work found that once a…
Macro-Operator Generation and Predicate Selection for TAMP Operator Learning
arXiv:2608.23629v1 Announce Type: cross Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning…
Identifying Latent Declarative Representations of Code for Assisting Repository Migration
arXiv:2608.23619v1 Announce Type: cross Abstract: Legacy software repositories embed decades of domain knowledge in undocumented code, making…
SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
arXiv:2608.24870v1 Announce Type: new Abstract: Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for…
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
arXiv:2608.24876v1 Announce Type: new Abstract: Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the…
