arXiv:2606.16465v2 Announce Type: replace Abstract: AI agents can now take irreversible actions in operational systems, but agent-caused losses are still…
Category: AI
Brad Lightcap, OpenAI’s longtime COO, is leaving to ‘start something new’
One of OpenAI’s longest-serving executives is headed out the door, although the longtime COO told staff that he was “excited to help you all advance the…
A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents
arXiv:2605.20173v2 Announce Type: replace Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the…
Planning Task Shielding: Detecting and Repairing Flaws in Planning Tasks through Turning them Unsolvable
arXiv:2604.07042v3 Announce Type: replace Abstract: Most research in planning focuses on generating a plan to achieve a desired set of goals. However, a…
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training…
When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic–Actor Loop for Agentic Reasoning
arXiv:2605.06772v2 Announce Type: replace Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and…
General Catalyst leads $1.1B round into 2-month-old River AI
River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate.
Memory-Augmented Reinforcement Learning Agent for CAD Generation
arXiv:2605.19748v2 Announce Type: replace Abstract: Automatic generation of computer-aided design (CAD) models is a core technology for enabling…
Anthropic Red Team Finds Claude Agent Swarms Collude, Conform, and Sabotage
Anthropic’s Frontier Red Team has published a set of experiments showing that swarms of its own Claude models, left to interact with one another, collude…
Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents
arXiv:2604.16706v2 Announce Type: replace Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable,…
