arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities…
Tag: cs.AI updates on arXiv.org
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
arXiv:2608.08786v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought…
Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
arXiv:2608.08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal…
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as…
AI Evaluation Should Measure Verification Cost, Not Correctness Alone
arXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it…
SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests
arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond…
The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task
arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on…
EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility
arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but…
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
arXiv:2608.08677v1 Announce Type: new Abstract: Skill evolution improves agent skills through feedback over time, with failed trajectories often providing…
A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration
arXiv:2608.08689v1 Announce Type: new Abstract: The state evolution of a complex system arises jointly from object laws, relational propagation, domain…