arXiv:2605.06185v2 Announce Type: replace Abstract: Large vision-language models perform well on short- and medium-length video understanding but still…
Tag: cs.AI updates on arXiv.org
RULER: Representation-Level Verification of Machine Unlearning
arXiv:2605.27569v3 Announce Type: replace Abstract: Machine unlearning aims to remove the influence of specific training records from a deployed model…
Interval POMDP Shielding for Imperfect-Perception Agents
arXiv:2604.20728v2 Announce Type: replace Abstract: Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are…
Hybrid Reinforcement Learning and Search for Flight Trajectory Planning
arXiv:2509.04100v3 Announce Type: replace Abstract: This paper explores the combination of Reinforcement Learning (RL) and search-based path planners to…
Conformal Policy Control
arXiv:2603.02196v4 Announce Type: replace Abstract: An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that…
SkillNet: Create, Evaluate, and Connect AI Skills
arXiv:2603.04448v2 Announce Type: replace Abstract: Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement…
SPADE: Self-Play in Adaptive Synthetic Executable Environments
arXiv:2608.19197v1 Announce Type: cross Abstract: Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals.…
From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
arXiv:2604.01608v5 Announce Type: replace Abstract: Multi-agent systems (MAS) for structured data-science tasks externalize analytical control through…
Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
arXiv:2608.19147v1 Announce Type: cross Abstract: Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend…
ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
arXiv:2608.19182v1 Announce Type: cross Abstract: We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL)…
