Anthropic’s Fable 5 is considered the most powerful AI model on the market, but U.S. companies are barely buying it. According to Ramp data, Fable 5…
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We…
AI News Brief Hourly Summary 2026-08-13 13h : 11 posts
11 posts were published in the last hour 10:32 : Claim-Level Reliability Assessment for Efficient Test-Time Reasoning 10:32 : Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges 10:32 : Mechanist: AI as a Scientific Instrument for Discovering…
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
arXiv:2608.11994v1 Announce Type: new Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through…
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
arXiv:2608.12097v1 Announce Type: new Abstract: Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to…
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their…
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
arXiv:2608.12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must…
Anthropic brings Claude Cowork to its Chrome extension, adding skills and plugins to the browser
Claude Cowork now runs directly in the side panel of Anthropic’s Chrome extension. The article Anthropic brings Claude Cowork to its Chrome extension,…
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
arXiv:2608.11977v1 Announce Type: new Abstract: Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed…
OEIS Open: How many conjectures can language models turn into theorems?
arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized…
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
arXiv:2608.11775v1 Announce Type: new Abstract: Gist-based context compression—summarising older conversation history into compact representations—is…
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
arXiv:2608.11949v1 Announce Type: new Abstract: Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent…
Policy-as-logic for robust reasoning over rules
arXiv:2608.11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance,…
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape…
AI News Brief Hourly Summary 2026-08-13 12h : 13 posts
13 posts were published in the last hour 9:32 : Proportional Analogies on Probability Distributions via Bayesian Updating 9:32 : HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting 9:32 : Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents…
Proportional Analogies on Probability Distributions via Bayesian Updating
arXiv:2608.11724v1 Announce Type: new Abstract: Analogies are quaternary relations of the form “A is to B as C is to D”. Among the various formalizations…
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
arXiv:2608.11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which…
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
arXiv:2608.11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing…
