arXiv:2609.04172v1 Announce Type: new Abstract: On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a…
Tag: AI
A Computationally Feasible Framework for Causal Probabilistic Explanation
arXiv:2609.04177v1 Announce Type: new Abstract: Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to…
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
arXiv:2609.04198v1 Announce Type: new Abstract: Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then…
Nvidia’s AI Sector Equity Stakes Grow, SEC Filing Shows
Nvidia’s equity investments across the artificial intelligence sector have grown sharply over the past year, according to the company’s quarterly filing…
A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
arXiv:2609.04170v1 Announce Type: new Abstract: Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate,…
The Natural Language Interaction Protocol and Standard for AI Agents
arXiv:2609.04135v1 Announce Type: new Abstract: AI agents are increasingly being developed and deployed across organizations using heterogeneous…
Efficient Test-Time Adaptation through Human-AI Interaction
arXiv:2609.04141v1 Announce Type: new Abstract: AI agents are trained on population-scale data to encode broad capabilities spanning those of many…
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
arXiv:2609.04148v1 Announce Type: new Abstract: As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while…
M&T Bank expands enterprise AI after years of technology overhaul
M&T Bank has deployed AI copilots to more than 15,000 employees as the US regional bank applies AI to internal operations, customer service, software…
Environment Evolution for Terminal Agents
arXiv:2609.04128v1 Announce Type: new Abstract: Scaling interactive and verifiable environments is critical for training terminal agents. As frontier…
