VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

arXiv:2610.10782v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many
become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the
visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor
and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as
scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty
is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with
actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and
visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest
self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized
RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution,
VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: