arXiv:2606.06081v2 Announce Type: replace Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.…
Category: cs.AI updates on arXiv.org
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
arXiv:2606.21654v2 Announce Type: replace Abstract: Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop…
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
arXiv:2606.16149v4 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer;…
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
arXiv:2605.02782v2 Announce Type: replace Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech.…
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
arXiv:2605.22664v5 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts…
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
arXiv:2605.06185v2 Announce Type: replace Abstract: Large vision-language models perform well on short- and medium-length video understanding but still…
RULER: Representation-Level Verification of Machine Unlearning
arXiv:2605.27569v3 Announce Type: replace Abstract: Machine unlearning aims to remove the influence of specific training records from a deployed model…
Interval POMDP Shielding for Imperfect-Perception Agents
arXiv:2604.20728v2 Announce Type: replace Abstract: Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are…
Hybrid Reinforcement Learning and Search for Flight Trajectory Planning
arXiv:2509.04100v3 Announce Type: replace Abstract: This paper explores the combination of Reinforcement Learning (RL) and search-based path planners to…
Conformal Policy Control
arXiv:2603.02196v4 Announce Type: replace Abstract: An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that…
