Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

arXiv:2602.02043v2 Announce Type: replace-cross
Abstract: Modern vision-language models struggle with basic compositional reasoning, failing to bind attributes to objects or relations to their referents. Existing benchmarks either rely on noisy real images that conflate confounding visual variables with the reasoning failure, or use simplistic synthetic scenes lacking the realism modern VLMs are tuned for. We introduce \textbf{Auto-Comp}, a fully automated, concept-driven pipeline that bridges this gap by generating photorealistic compositional benchmarks at scale. Its core innovation is a \textit{parallel A/B construction}: for each concept, the pipeline emits a \textit{Minimal} sample (template caption, isolated objects on a white background) and a \textit{Contextual} sample (LLM-rewritten caption, objects embedded in a realistic scene), isolating core binding ability from visio-linguistic complexity. We instantiate \textit{four} task families spanning the two canonical axes of compositional binding: \textit{Color} and \textit{Shape-Color} (attribute binding), and \textit{Position} and \textit{Relative Size} (relational binding). We evaluate over 25 VLMs spanning CLIP, SigLIP, hard-negative-trained, and frontier generative models. The findings are consistent across architectures and scales: every model exhibits a large Swap-vs-Confusion gap, with low-entropy distractors (e.g., repeated objects or colors) exposing failures \textit{beyond} the known bag-of-words limitations. We further uncover a task-dependent trade-off: visio-linguistic context aids relational reasoning but hinders attribute binding through visual clutter. We publicly release the pipeline and benchmarks.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: