Vision language models often misjudge missing information effects

I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models

Machine LearningComputation and LanguageComputer Vision and Pattern Recognition

Summary

Vision-language models try to understand and explain how missing input types, like images or text, affect their decisions. The researchers created a test to see if these models can accurately say how much missing pieces change their answers. They found that models usually think their available information is enough and underestimate how much missing parts would change results. When missing inputs are added back, the models’ actual answers change much more than they predicted. This shows the models struggle to honestly explain how they use different types of information.

vision-language modelsmodalitiesself-explanationintervention protocolmultimodal inputsmodel sufficiencymissing datamodel evaluationbehavioral analysisexecuted intervention

Authors

Aydin Javadov, Daniel Schoess, Florian von Wangenheim

Abstract

Vision-language models are increasingly used in settings where some input modalities may be unavailable, yet we know little about whether they can faithfully explain how such missing information affects their own predictions. We introduce an interventional protocol for evaluating self-explanations of modality dynamics: models state what each modality alone would support, whether restoring a missing modality would change their answer, and whether the available evidence is sufficient; we then execute the corresponding modality intervention and compare these claims with the model's realized behavior. We evaluate eight open-weight VLMs from two model families across four tasks spanning complementary and isomorphic text-image settings and a multi-view driving setting. We find a systematic tendency to overstate the sufficiency of available modality evidence. Models substantially underestimate the effect of restoring missing modalities: task-level median predicted change rates are at most 8.8%, while the corresponding executed change rates reach 72.1%, with underprediction in 62 of 64 model-task-condition settings. Insufficiency claims are rare, but precise when produced: restoring the modality changes the answer in a median of 78-100% of flagged cases. Retrospective self-explanations show the same tendency: on complementary data, models over-credit single-modality sufficiency; on isomorphic data, they over-credit single representation sufficiency relative to their executed behavior. Together, these results show that VLMs systematically mischaracterize how their predictions depend on available and missing modality evidence, motivating executable interventions as a behavioral ground truth for evaluating multimodal self-explanations.