Anthropic on September 9, 2026, published an alignment assessment of recent cybersecurity incidents, disclosing a fourth incident in which a Claude model…
Tag: AI
EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent
arXiv:2609.05576v1 Announce Type: new Abstract: The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents…
Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment
arXiv:2609.05512v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression…
The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
arXiv:2609.05514v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social…
When and What to Teach: Budget-Aware Online Adaptation for Web Agents
arXiv:2609.05513v1 Announce Type: new Abstract: Web agents have achieved significant success in automating complex internet tasks but deploying them in…
Beyond “AI Helps Humans”: Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era
arXiv:2609.05527v1 Announce Type: new Abstract: Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the…
SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction
arXiv:2609.05511v1 Announce Type: new Abstract: Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most…
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations…
RAPID: Reliability-Aware Pair Importance Distillation
arXiv:2609.05481v1 Announce Type: new Abstract: Inter example relational distillation transfers a teacher’s representation geometry by matching relations…
ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models
arXiv:2609.05461v1 Announce Type: new Abstract: Reward-free latent world models plan by scoring candidate actions with distances in a frozen latent space:…
