arXiv:2606.31672v5 Announce Type: replace-cross Abstract: Despite rapid progress in interactive world models (IWMs), short-horizon performance does not…
Tag: AI
AgenticRL: Agentic Reinforcement Learning with Self-Refinement for Complex UAV Navigation
arXiv:2606.03963v4 Announce Type: replace-cross Abstract: Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but…
GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
arXiv:2607.03869v2 Announce Type: replace-cross Abstract: Referring remote sensing image segmentation segments the object named by a natural-language…
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
Many developers find that an agent idea works inside Claude Code or Codex, then struggles once they rebuild it with their own loop. The Strands Agents…
Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models
arXiv:2606.04287v2 Announce Type: replace-cross Abstract: Generating realistic and diverse graphs is a key problem in machine learning, with applications…
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation
arXiv:2605.06445v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation…
PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
arXiv:2606.00515v2 Announce Type: replace-cross Abstract: Contact-rich manipulation demands both high-level semantic reasoning and the safe regulation of…
The critical slowing down in training diffusion models
arXiv:2605.12597v3 Announce Type: replace-cross Abstract: Computational sampling has been central to the sciences since the mid-20th century. While…
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
arXiv:2605.12969v5 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for…
LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
arXiv:2605.09384v2 Announce Type: replace-cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment…
