arXiv:2607.26643v2 Announce Type: replace Abstract: Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions…
Tag: cs.AI updates on arXiv.org
BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
arXiv:2608.04156v2 Announce Type: replace Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it…
Fragility of Value under Imperfect Alignment
arXiv:2607.28881v3 Announce Type: replace Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that…
ContextSniper: AntTrail’s Token-Efficient Code Memory for Repository-Level Program Repair
arXiv:2607.01916v5 Announce Type: replace Abstract: Large language model agents can repair real repository issues, but they often spend large context…
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
arXiv:2607.14178v3 Announce Type: replace Abstract: Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex…
A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
arXiv:2606.06081v2 Announce Type: replace Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.…
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
arXiv:2606.21654v2 Announce Type: replace Abstract: Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop…
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
arXiv:2606.16149v4 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer;…
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
arXiv:2605.02782v2 Announce Type: replace Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech.…
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
arXiv:2605.22664v5 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts…
