arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model…
Category: cs.AI updates on arXiv.org
Generative Embodied Multiple Behavior Control Systems for Human-like Agents
arXiv:2609.22691v1 Announce Type: new Abstract: An enduring and richly elaborated dichotomy in cognitive neuroscience is that of human behavior control…
Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review
arXiv:2609.22694v1 Announce Type: new Abstract: Public crash databases increasingly support automated safety analysis, but crash severity prediction…
Self-Organizing Agent Teams Learn to Reason Together
arXiv:2609.22682v1 Announce Type: new Abstract: Collective intelligence depends not only on what team members know, but also on how they organize their…
A Survey on the Linear Representation Hypothesis
arXiv:2609.22695v1 Announce Type: new Abstract: The term “linear representation hypothesis” (LRH) has appeared across diverse subfields of artificial…
AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
arXiv:2609.22592v1 Announce Type: new Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment…
Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
arXiv:2609.22620v1 Announce Type: new Abstract: Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split…
GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
arXiv:2609.22619v1 Announce Type: new Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation…
EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
arXiv:2609.22537v1 Announce Type: new Abstract: Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence.…
MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent…
