arXiv:2408.10608v2 Announce Type: replace-cross Abstract: Large language models (LLMs) may encode biased associations from heterogeneous training corpora…
Author: script
Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what…
Incentives to Offer Algorithmic Recourse
arXiv:2301.12884v2 Announce Type: replace-cross Abstract: Algorithmic recourse promises to help applicants rejected by automated systems by explaining the…
Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses
arXiv:2609.05736v2 Announce Type: replace Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed…
FrontierChallenge: Evaluating Scientific Workflow Completion
arXiv:2608.24979v2 Announce Type: replace Abstract: Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most…
One week left to book your exhibit table at TechCrunch Disrupt 2026
Only one week left to secure your exhibit table. Tables are limited and can sell out before the September 18 deadline.
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
arXiv:2609.01315v2 Announce Type: replace Abstract: Building an omni-modal foundation model means evaluating it across text, image, video, and audio.…
OpenAI’s feud with mathematicians is only escalating
Twenty-five leading mathematicians signed an open letter arguing that AI labs are threatening their intellectual work.
From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale
arXiv:2609.05758v2 Announce Type: replace Abstract: Conversational assistants can blend retrieval, action selection, escalation, and wording in a single…
Y Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, too
Tan argues that frontier models themselves trained on public human knowledge so access to capable AI should be “a form of public good.”
