arXiv:2608.11994v1 Announce Type: new Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through…
Author: script
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
arXiv:2608.12097v1 Announce Type: new Abstract: Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to…
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their…
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
arXiv:2608.12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must…
Anthropic brings Claude Cowork to its Chrome extension, adding skills and plugins to the browser
Claude Cowork now runs directly in the side panel of Anthropic’s Chrome extension. The article Anthropic brings Claude Cowork to its Chrome extension,…
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
arXiv:2608.11977v1 Announce Type: new Abstract: Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed…
OEIS Open: How many conjectures can language models turn into theorems?
arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized…
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
arXiv:2608.11775v1 Announce Type: new Abstract: Gist-based context compression—summarising older conversation history into compact representations—is…
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
arXiv:2608.11949v1 Announce Type: new Abstract: Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent…
Policy-as-logic for robust reasoning over rules
arXiv:2608.11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance,…
