Multimodal AI models overdescribe images compared to humans using culture
Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture
Computation and LanguageArtificial IntelligenceHuman-Computer Interaction
Summary
This paper looks at how people and AI systems describe images with clues in a game called Dixit. People naturally use hints that can be understood in different ways and often include cultural references. The authors found that AI models tend to give very specific descriptions that don't allow for multiple interpretations, and rarely mention cultural ideas. This shows that AI might miss an important part of how humans communicate when using images and words together.
What this means in practice
- •For multimodal ai developers: Improve the design of AI image captioning systems to produce clues that allow multiple valid interpretations and include cultural context.
- •For creative marketing teams: Create richer, culturally resonant image captions and marketing messages that mimic human ambiguity for engagement.
Authors
Cody Kommers, Mingrui Ye, Evelyn Gius, Daniela Mihai, Hoyt Long, Zheng Yuan, Drew Hemment
Abstract
Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. From humour to politics to art, people express themselves in words and images that are open enough to invite different interpretations, yet constrained enough to be interpretable. We operationalise this notion of calibrated ambiguity with a task drawn from the parlour game Dixit. We compare differences in clues generated by human vs multimodal language models, based on a novel coding rubric for calibrated ambiguity, and find that models consistently exhibit ambiguity collapse (i.e., their outputs are over-specified, leaving no room for multiple legitimate interpretations). Unlike human clues, AI-generated clues also exhibit cultural flattening; they almost never make reference to culturally-situated knowledge, even when prompted to use allusion and figurative language.