arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it…
Category: cs.AI updates on arXiv.org
Back to the Future: A workbook time machine for spread sheet creation benchmarks
arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the…
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage…
Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes
arXiv:2608.07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes…
AndroidReality: How Far Are Mobile Agents from the Real World?
arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their…
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization’s reputation and attract customers, but only when results are…
The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
arXiv:2608.07779v1 Announce Type: new Abstract: Artificial intelligence is changing the task composition of computing work faster than curricula and…
IntelliAudit: Using Large Language Models to Evaluate Audit Controls
arXiv:2608.07688v1 Announce Type: new Abstract: IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic…
Protecting patient privacy in clinical foundation models: Technical and legal perspectives
arXiv:2608.07705v1 Announce Type: new Abstract: Clinical foundation models trained on large-scale patient data are increasingly used for decision support,…
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination,…
