11 posts were published in the last hour
- 10:32 : Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
- 10:32 : Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
- 10:32 : Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
- 10:32 : CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
- 10:32 : Anthropic brings Claude Cowork to its Chrome extension, adding skills and plugins to the browser
- 10:32 : Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
- 10:3 : OEIS Open: How many conjectures can language models turn into theorems?
- 10:3 : The Sleeping Agent: What Gist-Based Context Compression Loses and Why
- 10:3 : ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
- 10:3 : Policy-as-logic for robust reasoning over rules
- 10:3 : Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents