arXiv:2609.09664v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as personalized assistants that interact with…
Tag: AI
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
arXiv:2609.11243v2 Announce Type: replace Abstract: Autonomous research agents are increasingly expected to search the literature, analyze experimental…
Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
arXiv:2609.15293v3 Announce Type: replace Abstract: Binding the audit flag in Reflexion-style agents — without changing the auditor — reduces attack…
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
arXiv:2608.03952v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to provide conversational practice for…
Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning
arXiv:2608.23982v2 Announce Type: replace Abstract: Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it…
HANIA: Planner-Guided Multimodal Graph Evidence Selection for Grounded Question Answering
arXiv:2608.29088v2 Announce Type: replace Abstract: Multimodal question answering remains sensitive to noisy, incomplete, and weakly grounded evidence.…
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
arXiv:2609.09418v3 Announce Type: replace Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing…
U.S. TRANSCOM deploys randomised AI to secure military logistics
Deploying randomised AI logistics offers military planners a viable defence against adversarial tracking, allowing U.S. Transportation Command (TRANSCOM)…
Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
arXiv:2607.11433v2 Announce Type: replace Abstract: Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer…
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
arXiv:2606.27161v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but…
