Vision language models often rely too much on language cues

When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models

Computation and Language

Summary

Vision-language models try to understand pictures and text together but sometimes pay too much attention to the words and ignore the images, leading to mistakes. The authors create a way to break down and track how this problem happens inside the models. They focus on two ideas: how much the model expects certain words (linguistic priors) and how well the words match the images (cross-modal coverage). By studying these through four steps, they reveal how language bias makes the models get things wrong.

What this means in practice

Authors

Yizhou Fang, Siyue Chen, Zimo Qi, Zhiyu Xue, Xi Chen, Guangliang Liu

Abstract

Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.