IBM announced a strategic partnership with OpenAI on August 13, 2026, that embeds OpenAI’s frontier models, including GPT-5.6, and products such as Codex…
Category: AI
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
arXiv:2608.12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets. External oracles…
Fable 5’s slow adoption suggests corporate willingness to pay for frontier AI has hit a ceiling
Anthropic’s Fable 5 is considered the most powerful AI model on the market, but U.S. companies are barely buying it. According to Ramp data, Fable 5…
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We…
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
arXiv:2608.11994v1 Announce Type: new Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through…
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
arXiv:2608.12097v1 Announce Type: new Abstract: Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to…
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their…
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
arXiv:2608.12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must…
Anthropic brings Claude Cowork to its Chrome extension, adding skills and plugins to the browser
Claude Cowork now runs directly in the side panel of Anthropic’s Chrome extension. The article Anthropic brings Claude Cowork to its Chrome extension,…
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
arXiv:2608.11977v1 Announce Type: new Abstract: Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed…
