arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment…
Tag: AI
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it…
How AI is changing the vulnerability response timeline
Artificial intelligence is giving security researchers new ways to examine code, trace unusual behaviour and identify flaws that conventional tools may…
Back to the Future: A workbook time machine for spread sheet creation benchmarks
arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the…
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage…
Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes
arXiv:2608.07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes…
AndroidReality: How Far Are Mobile Agents from the Real World?
arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their…
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization’s reputation and attract customers, but only when results are…
webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware
webAI has released TwIL-LM, a family of formal-logic models at 1.7B and 3B parameters that translate English into first-order logic and check whether…
The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
arXiv:2608.07779v1 Announce Type: new Abstract: Artificial intelligence is changing the task composition of computing work faster than curricula and…