(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

arXiv:2609.00443v1 Announce Type: cross
Abstract: Language models learn about grammatical number primarily from co-occurrence, and show frequency effects as a result—sometimes taken to indicate that they do not learn abstract “rules'', and are instead dependent on specific lexical items. Testing generalization with text stimuli alone cannot settle this debate, since distributional cues (is/are, this/these) easily give number away. We instead use cross-modal generalization as a tool to investigate abstractions in LMs that can also accept visual inputs (VLMs), restricting the evidence that diagnoses number to an extra-linguistic modality. We teach VLMs pairs of new nouns by adding new embeddings and only updating them during learning, comparing conditions where number is diagnosed by visual cues alone against ones where it is disambiguated by text. Across behavior, representational dynamics, and causal mechanisms, we find non-trivial evidence for cross-modal generalization across both exposure conditions, and that linguistic vs. extra-linguistic cue conditions are treated in similar ways in the internal mechanisms of the model. This suggests that statistical learners like VLMs can generalize beyond surface-level co-occurrence and show genuine abstraction-compatible behavior.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: