arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment…
Tag: cs.AI updates on arXiv.org
LLM Agents for Time-Series: A Survey
arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary…
Same Model, Different Harness: Different Coding-Agent Results
arXiv:2608.26218v1 Announce Type: new Abstract: A coding agent combines a model with a harness, which decides what the model sees, which tools it can use,…
Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs
arXiv:2608.26191v1 Announce Type: new Abstract: Incident risk prediction from longitudinal electronic health records (EHRs) is challenging because…
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
arXiv:2608.26199v1 Announce Type: new Abstract: We ask whether AI agents powered by locally deployed large language models can reliably automate…
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
arXiv:2608.26190v1 Announce Type: new Abstract: World models enable agents to reason about future outcomes and learn policies from their knowledge of…
Agentic AI for operating scientific instruments for nanoscale characterization
arXiv:2608.26198v1 Announce Type: new Abstract: Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert…
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
arXiv:2608.26193v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet…
Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion
arXiv:2608.26185v1 Announce Type: new Abstract: Equal participation in co-located discussion is important for effective collaboration, yet people often…
TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education
arXiv:2608.26184v1 Announce Type: new Abstract: AI programming tutors provide scalable support, yet lack the behavioral context human tutors rely on to…
