arXiv:2609.17755v1 Announce Type: cross Abstract: The analysis of US political language is usually based on the written form (e.g. presidential addresses)…
Tag: cs.AI updates on arXiv.org
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
arXiv:2609.17708v1 Announce Type: cross Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models:…
One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG
arXiv:2609.17709v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator…
CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors
arXiv:2609.17758v1 Announce Type: cross Abstract: Deep Reinforcement Learning has demonstrated remarkable capability in quadrotor control, yet learned…
REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
arXiv:2609.17745v1 Announce Type: cross Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external…
Accelerating Diffusion Sampling via Speculative Draft Trees
arXiv:2609.17691v1 Announce Type: cross Abstract: Speculative sampling accelerates diffusion model generation by drafting inexpensive candidate states and…
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
arXiv:2609.17653v1 Announce Type: cross Abstract: GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed…
Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models
arXiv:2609.17644v1 Announce Type: cross Abstract: Domain-specialized language models are widely used for scientific question answering, but stronger…
Scaling Articulated Rationales for MLLM-based Recommendation
arXiv:2609.17639v1 Announce Type: cross Abstract: Modern recommendation systems largely infer user preferences from implicit behaviors such as clicks,…
The Missing “I Don’t Know”: Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention
arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show…
