Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

Computation and Language

Summary

The authors looked at how humans and vision-language AI models understand descriptions in repeated naming games where people describe new things. They found that humans use context well to understand what others mean, even as the conversation changes. The AI models could somewhat use past context but had trouble keeping track of important details over time to understand the descriptions fully. This shows current models still miss key abilities needed for smooth back-and-forth communication.

Authors

Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce

Abstract

Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.