arXiv:2608.11698v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level…
Category: AI
Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning
arXiv:2608.11658v1 Announce Type: cross Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an…
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
arXiv:2608.11674v1 Announce Type: cross Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they…
Amazon Quick Arrives Inside Word, Excel, PowerPoint, and Outlook
AWS has brought its Amazon Quick assistant directly into Microsoft 365, announcing on August 13, 2026 that extensions for Word, Excel, PowerPoint, and…
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
arXiv:2608.11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are…
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output tokens per…
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
arXiv:2608.11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard…
What We Learned by Reproducing 2,200 papers from ICML
This post has no text preview — click the link below to read the original article. This article has been indexed from Hugging Face – Blog Read the original article: What We Learned by Reproducing 2,200 papers from ICML
Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
arXiv:2608.11661v1 Announce Type: cross Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner…
Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones
arXiv:2608.11627v1 Announce Type: cross Abstract: The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer…
