arXiv:2608.08802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more…
Tag: cs.AI updates on arXiv.org
Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models
arXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet…
Three Generations of Healthcare IT: From the Digital Record to the Computable Care Process
arXiv:2608.08806v1 Announce Type: new Abstract: Objective. Healthcare IT is usually organized by the technologies it adopts. We instead organize it by the…
PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary
arXiv:2608.08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual…
Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs
arXiv:2608.08794v1 Announce Type: new Abstract: Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial…
FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities…
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
arXiv:2608.08786v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought…
Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
arXiv:2608.08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal…
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as…
AI Evaluation Should Measure Verification Cost, Not Correctness Alone
arXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it…