arXiv:2609.02057v1 Announce Type: new Abstract: Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits…
Category: AI
Anthropic’s new Fable release is cheaper, less restrictive
Fable 5.1 includes changes meant to reduce token cost and false-positive restrictions from the model’s safeguards.
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from…
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
arXiv:2609.01924v1 Announce Type: new Abstract: Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard…
When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor
arXiv:2609.01985v1 Announce Type: new Abstract: As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization…
Benchmarking Language Models for Statistical Problem Formulation
arXiv:2609.01982v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as assistants for statistical and data science work,…
The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction
arXiv:2609.01909v1 Announce Type: new Abstract: Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available…
Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment
arXiv:2609.01962v1 Announce Type: new Abstract: Ultra-low-bit language models can reduce storage and memory bandwidth, but a nominal “1.58-bit” label does…
Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
arXiv:2609.01873v1 Announce Type: new Abstract: Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is…
Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization
arXiv:2609.01861v1 Announce Type: new Abstract: The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve…
