arXiv:2606.31672v5 Announce Type: replace-cross Abstract: Despite rapid progress in interactive world models (IWMs), short-horizon performance does not…
Tag: cs.AI updates on arXiv.org
AgenticRL: Agentic Reinforcement Learning with Self-Refinement for Complex UAV Navigation
arXiv:2606.03963v4 Announce Type: replace-cross Abstract: Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but…
GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
arXiv:2607.03869v2 Announce Type: replace-cross Abstract: Referring remote sensing image segmentation segments the object named by a natural-language…
Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models
arXiv:2606.04287v2 Announce Type: replace-cross Abstract: Generating realistic and diverse graphs is a key problem in machine learning, with applications…
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation
arXiv:2605.06445v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation…
PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
arXiv:2606.00515v2 Announce Type: replace-cross Abstract: Contact-rich manipulation demands both high-level semantic reasoning and the safe regulation of…
The critical slowing down in training diffusion models
arXiv:2605.12597v3 Announce Type: replace-cross Abstract: Computational sampling has been central to the sciences since the mid-20th century. While…
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
arXiv:2605.12969v5 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for…
LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
arXiv:2605.09384v2 Announce Type: replace-cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment…
REALM: An RGB- and Event-Aligned Latent Manifold for Cross-Modal Perception
arXiv:2605.00271v4 Announce Type: replace-cross Abstract: Event cameras provide several unique advantages over standard frame-based sensors, including…
