Vision language models often rely too much on language cues
When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models
Computation and Language
Summary
Vision-language models try to understand pictures and text together but sometimes pay too much attention to the words and ignore the images, leading to mistakes. The authors create a way to break down and track how this problem happens inside the models. They focus on two ideas: how much the model expects certain words (linguistic priors) and how well the words match the images (cross-modal coverage). By studying these through four steps, they reveal how language bias makes the models get things wrong.
What this means in practice
- •For machine learning engineers: Diagnose and reduce language bias in vision-language models by tracing bias propagation through inference steps.
- •For computer vision teams: Improve model accuracy by balancing linguistic cues and visual content in multimodal applications.
Authors
Yizhou Fang, Siyue Chen, Zimo Qi, Zhiyu Xue, Xi Chen, Guangliang Liu
Abstract
Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.