arXiv:2312.17295v2 Announce Type: replace-cross Abstract: With the rise of large language models (LLMs) and concerns about potential misuse, watermarks…
Tag: AI
WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
arXiv:2609.23806v2 Announce Type: replace Abstract: Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for…
A Scalable Multi-Robot Framework for Decentralized and Asynchronous Perception-Action-Communication Loops
arXiv:2309.10164v3 Announce Type: replace-cross Abstract: We develop a decentralized Perception-Action-Communication (PAC) system for multi-robot teams…
Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models
arXiv:2609.21113v2 Announce Type: replace Abstract: Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream…
Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving
arXiv:2609.21470v2 Announce Type: replace Abstract: Conventional end-to-end driving systems model the environment with sparse objects and lane elements.…
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
arXiv:2609.21423v2 Announce Type: replace Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and…
A Lie Detector Test for Language Models: Reading Knowledge a Model Won’t Reveal
arXiv:2609.21996v2 Announce Type: replace Abstract: Large language models can hold knowledge they do not report. A model may sandbag on a capability…
QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training
arXiv:2609.19513v2 Announce Type: replace Abstract: High-quality pre-training data is a critical bottleneck for educational and STEM-specific language…
From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development
arXiv:2609.11493v2 Announce Type: replace Abstract: Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of…
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
arXiv:2609.17885v2 Announce Type: replace Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet…
