arXiv:2607.19364v2 Announce Type: replace Abstract: Activation steering adds a residual-stream direction at inference time, providing lightweight…
Tag: cs.AI updates on arXiv.org
UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention
arXiv:2607.17188v2 Announce Type: replace Abstract: While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through…
AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning
arXiv:2606.00671v2 Announce Type: replace Abstract: We present AXIOM, a trust-first neuro-symbolic architecture for natural-language mathematical…
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
arXiv:2605.27944v2 Announce Type: replace Abstract: With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly…
A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline
arXiv:2606.07718v2 Announce Type: replace Abstract: Agentic AI offers a promising path to automating software development bottlenecks in scientific…
When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting
arXiv:2606.16465v2 Announce Type: replace Abstract: AI agents can now take irreversible actions in operational systems, but agent-caused losses are still…
A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents
arXiv:2605.20173v2 Announce Type: replace Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the…
Planning Task Shielding: Detecting and Repairing Flaws in Planning Tasks through Turning them Unsolvable
arXiv:2604.07042v3 Announce Type: replace Abstract: Most research in planning focuses on generating a plan to achieve a desired set of goals. However, a…
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training…
When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic–Actor Loop for Agentic Reasoning
arXiv:2605.06772v2 Announce Type: replace Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and…
