arXiv:2609.05251v1 Announce Type: new Abstract: We present a unified reinforcement-learning (RL) framework that discovers compact parametrized quantum…
Tag: cs.AI updates on arXiv.org
Uncensored Open-weight Models: Redistribution as the Persistence Layer
arXiv:2609.05241v1 Announce Type: new Abstract: A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models.…
Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents
arXiv:2609.05261v1 Announce Type: new Abstract: Large language model agents increasingly rely on execution traces to master complex interactive tasks.…
Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory
arXiv:2609.05245v1 Announce Type: new Abstract: Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery…
Substrate-Aware AI Agents: Execution Context as a First-Class Input
arXiv:2609.05232v1 Announce Type: new Abstract: Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime,…
The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior
arXiv:2609.05190v1 Announce Type: new Abstract: In this paper we illustrate a novel architecture generating interpretable behavior and explanations. We…
CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review
arXiv:2609.05227v1 Announce Type: new Abstract: Recent reports during the AAAI-27 review cycle highlight the risk of reviewers coordinating bids for…
ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs
arXiv:2609.05228v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models…
What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection
arXiv:2609.05198v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has emerged as a widely adopted post-training paradigm for enhancing large…
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
arXiv:2609.05141v1 Announce Type: new Abstract: Scientific papers require models to reason jointly over text, equations, figures, tables, code, and…
