arXiv:2609.15471v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a central approach for improving…
Category: cs.AI updates on arXiv.org
Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models
arXiv:2609.15517v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and…
Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents
arXiv:2609.15422v1 Announce Type: new Abstract: AI agents are provisioned the same as employee-owned hosts in many enterprise settings with a static…
Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA
arXiv:2609.15530v1 Announce Type: new Abstract: We describe our submission to the MedReason 2026 challenge, covering multiple-choice (MCQ) and open-ended…
SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution
arXiv:2609.15396v1 Announce Type: new Abstract: LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt…
Can AI systems have free will?
arXiv:2609.15407v1 Announce Type: new Abstract: While there has been much discussion of whether AI systems could function as moral agents or acquire…
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary
arXiv:2609.15397v1 Announce Type: new Abstract: AI agents increasingly execute long-running workflows that externalize effects through independently…
Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning
arXiv:2609.15404v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (OPD) is becoming the standard way to integrate specialist…
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
arXiv:2609.15364v1 Announce Type: new Abstract: Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not…
Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
arXiv:2609.15293v1 Announce Type: new Abstract: When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results…
