arXiv:2609.21113v2 Announce Type: replace Abstract: Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream…
Tag: cs.AI updates on arXiv.org
Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving
arXiv:2609.21470v2 Announce Type: replace Abstract: Conventional end-to-end driving systems model the environment with sparse objects and lane elements.…
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
arXiv:2609.21423v2 Announce Type: replace Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and…
A Lie Detector Test for Language Models: Reading Knowledge a Model Won’t Reveal
arXiv:2609.21996v2 Announce Type: replace Abstract: Large language models can hold knowledge they do not report. A model may sandbag on a capability…
QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training
arXiv:2609.19513v2 Announce Type: replace Abstract: High-quality pre-training data is a critical bottleneck for educational and STEM-specific language…
From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development
arXiv:2609.11493v2 Announce Type: replace Abstract: Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of…
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
arXiv:2609.17885v2 Announce Type: replace Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet…
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations
arXiv:2609.09664v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as personalized assistants that interact with…
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
arXiv:2609.11243v2 Announce Type: replace Abstract: Autonomous research agents are increasingly expected to search the literature, analyze experimental…
Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
arXiv:2609.15293v3 Announce Type: replace Abstract: Binding the audit flag in Reflexion-style agents — without changing the auditor — reduces attack…
