arXiv:2609.04141v1 Announce Type: new Abstract: AI agents are trained on population-scale data to encode broad capabilities spanning those of many…
Category: cs.AI updates on arXiv.org
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
arXiv:2609.04148v1 Announce Type: new Abstract: As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while…
Environment Evolution for Terminal Agents
arXiv:2609.04128v1 Announce Type: new Abstract: Scaling interactive and verifiable environments is critical for training terminal agents. As frontier…
From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
arXiv:2609.04166v1 Announce Type: new Abstract: Research and news coverage of language-model deception increasingly attributes human-like mental-state…
IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations
arXiv:2609.04030v1 Announce Type: new Abstract: IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific…
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
arXiv:2609.04094v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most…
Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable
arXiv:2609.04127v1 Announce Type: new Abstract: Large language models are increasingly used to support organizational decisions, yet users often lack a…
Spurious Advantage Hidden in GRPO
arXiv:2609.04063v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable…
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
arXiv:2609.04098v1 Announce Type: new Abstract: Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose…
LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
arXiv:2609.04013v1 Announce Type: new Abstract: Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine…
