arXiv:2608.24764v1 Announce Type: new Abstract: Large language model agents are moving beyond conventional retrieval-augmented generation toward direct…
Category: cs.AI updates on arXiv.org
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
arXiv:2608.24777v1 Announce Type: new Abstract: LLM-based agents can interact with external environments through tool invocation, but this capability also…
Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
arXiv:2608.24790v1 Announce Type: new Abstract: Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the…
Meta$^n$: Recursive Self-Improvement through Emergent Depth
arXiv:2608.24735v1 Announce Type: new Abstract: Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a…
Lifted Model Construction under Approximate Commutativity
arXiv:2608.24713v1 Announce Type: new Abstract: Lifted inference algorithms enable scalable probabilistic inference even for large object domains by…
Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
arXiv:2608.24658v1 Announce Type: new Abstract: Scaling test-time reasoning has substantially improved the problem-solving ability of large language…
Confident at the moment of action: belief miscalibration in LLM play under hidden information
arXiv:2608.24691v1 Announce Type: new Abstract: Agentic systems increasingly gate actions on a model’s own stated confidence, which assumes confidence…
The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
arXiv:2608.24662v1 Announce Type: new Abstract: Large language models (LLMs) are commonly evaluated under the assumption that their observable behavior is…
Joint Optimization of Tool Creation and Use for Large Language Model Agents
arXiv:2608.24571v1 Announce Type: new Abstract: Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation…
Causal Modelling of Support Interventions for Student Competency Assessment
arXiv:2608.24632v1 Announce Type: new Abstract: Accurate assessment of student competencies is essential for enabling educators to identify individual…
