Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation
Computation and Language
Summary
The authors looked at how humans and vision-language AI models understand descriptions in repeated naming games where people describe new things. They found that humans use context well to understand what others mean, even as the conversation changes. The AI models could somewhat use past context but had trouble keeping track of important details over time to understand the descriptions fully. This shows current models still miss key abilities needed for smooth back-and-forth communication.
iterated reference gamespragmatic reasoningcontext sensitivityvision-language modelslinguistic collaborationmulti-turn dialoguereferential expressionscontext building
Authors
Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce
Abstract
Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.