arXiv:2609.21470v2 Announce Type: replace Abstract: Conventional end-to-end driving systems model the environment with sparse objects and lane elements.…
Author: script
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
arXiv:2609.21423v2 Announce Type: replace Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and…
A Lie Detector Test for Language Models: Reading Knowledge a Model Won’t Reveal
arXiv:2609.21996v2 Announce Type: replace Abstract: Large language models can hold knowledge they do not report. A model may sandbag on a capability…
QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training
arXiv:2609.19513v2 Announce Type: replace Abstract: High-quality pre-training data is a critical bottleneck for educational and STEM-specific language…
AI News Brief Hourly Summary 2026-09-25 04h : 12 posts
12 posts published in the last hour 01:32From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development 01:32ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software 01:32PRAGMA: Evaluating Personalized Guidance with Memory Alignment in…
From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development
arXiv:2609.11493v2 Announce Type: replace Abstract: Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of…
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
arXiv:2609.17885v2 Announce Type: replace Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet…
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations
arXiv:2609.09664v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as personalized assistants that interact with…
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
arXiv:2609.11243v2 Announce Type: replace Abstract: Autonomous research agents are increasingly expected to search the literature, analyze experimental…
Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
arXiv:2609.15293v3 Announce Type: replace Abstract: Binding the audit flag in Reflexion-style agents — without changing the auditor — reduces attack…
