arXiv:2609.06079v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide…
Tag: cs.AI updates on arXiv.org
AutoResearch: Insight In, Hallucination Out
arXiv:2608.17906v4 Announce Type: replace Abstract: Autonomous research systems are increasingly capable of executing long research workflows, yet…
FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
arXiv:2608.08736v2 Announce Type: replace Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the…
AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance
arXiv:2608.21363v2 Announce Type: replace Abstract: Runtime-governance evidence often collapses materially different events into one audit record: a…
Faithful, Not Corrective: Model Capability Governs Message-Format Effects in Multi-Hop Agent Relays
arXiv:2607.09678v2 Announce Type: replace Abstract: When LLM agents hand information to one another, does the message format matter? Two literatures…
Multi-Resolution Attribution from Adaptive Routing State
arXiv:2605.22866v2 Announce Type: replace Abstract: Adaptive hierarchical systems accumulate routing state as they learn which components to select. We…
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
arXiv:2606.29251v3 Announce Type: replace Abstract: Financial decision-makers face more information than they can directly inspect, making context…
AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes
arXiv:2606.29871v2 Announce Type: replace Abstract: We present the AI Training Manager, a bounded LLM-based metacognitive monitoring-and-control layer for…
When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
arXiv:2608.05810v2 Announce Type: replace Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution…
A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
arXiv:2603.02540v2 Announce Type: replace Abstract: Large language models (LLMs) display a unified “general factor” of capability across 10 benchmarks (a…
