arXiv:2608.08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual…
Tag: cs.AI updates on arXiv.org
Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs
arXiv:2608.08794v1 Announce Type: new Abstract: Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial…
FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities…
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
arXiv:2608.08786v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought…
Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
arXiv:2608.08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal…
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as…
AI Evaluation Should Measure Verification Cost, Not Correctness Alone
arXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it…
SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests
arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond…
The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task
arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on…
EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility
arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but…
