arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided…
Tag: cs.AI updates on arXiv.org
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
arXiv:2608.07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer…
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment…
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it…
Back to the Future: A workbook time machine for spread sheet creation benchmarks
arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the…
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage…
Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes
arXiv:2608.07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes…
AndroidReality: How Far Are Mobile Agents from the Real World?
arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their…
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization’s reputation and attract customers, but only when results are…
The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
arXiv:2608.07779v1 Announce Type: new Abstract: Artificial intelligence is changing the task composition of computing work faster than curricula and…