arXiv:2607.08837v4 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the…
Category: AI
Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
arXiv:2607.13571v2 Announce Type: replace-cross Abstract: Sound event detection relies on frame-level strong labels whose annotation is expensive. Active…
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
arXiv:2607.21971v2 Announce Type: replace-cross Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by…
Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement
arXiv:2607.04277v2 Announce Type: replace-cross Abstract: The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement…
Git-Assistant: Planning-Based Support for Updating Git Repositories
arXiv:2607.09224v3 Announce Type: replace-cross Abstract: Version control systems are essential for collaborative software development, yet tools like git…
Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications
arXiv:2607.00442v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted,…
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
arXiv:2606.31672v5 Announce Type: replace-cross Abstract: Despite rapid progress in interactive world models (IWMs), short-horizon performance does not…
AgenticRL: Agentic Reinforcement Learning with Self-Refinement for Complex UAV Navigation
arXiv:2606.03963v4 Announce Type: replace-cross Abstract: Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but…
GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
arXiv:2607.03869v2 Announce Type: replace-cross Abstract: Referring remote sensing image segmentation segments the object named by a natural-language…
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
Many developers find that an agent idea works inside Claude Code or Codex, then struggles once they rebuild it with their own loop. The Strands Agents…
