arXiv:2604.10441v2 Announce Type: replace Abstract: Medical large language models are typically evaluated on idealized patient cases that do not reflect…
Tag: cs.AI updates on arXiv.org
TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas
arXiv:2603.16448v3 Announce Type: replace Abstract: Text-to-SQL parsing has achieved remarkable progress under the Full Schema Assumption. However, this…
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
arXiv:2604.06015v2 Announce Type: replace Abstract: Instruction tuning is commonly assumed to endow language models with a domain-general ability to…
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
arXiv:2603.15976v2 Announce Type: replace Abstract: While LLMs have accelerated scientific code generation, comprehensively evaluating generated code…
Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage
arXiv:2504.20007v4 Announce Type: replace Abstract: This paper proposes a novel interdisciplinary framework for analyzing police body-worn camera (BWC)…
GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay
arXiv:2609.11923v1 Announce Type: cross Abstract: Counterfactual regret minimization (CFR) is one of the few large numerical workloads that still runs…
Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning
arXiv:2506.14045v2 Announce Type: replace Abstract: Developing agents capable of exploring, planning and learning in complex open-ended environments is a…
Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
arXiv:2601.15397v3 Announce Type: replace Abstract: The rapid emergence of new entities — driven by cultural shifts, evolving trends, and personalized…
Timely Clinical Diagnosis through Active Test Selection
arXiv:2510.18988v5 Announce Type: replace Abstract: There is growing interest in using machine learning (ML) to support clinical diagnosis, but most…
Domain-Specific Hallucination Detection in Large Language Models
arXiv:2609.11878v1 Announce Type: cross Abstract: Large language models generate fluent text that can contain unfaithful claims — a phenomenon known as…
