Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts
2026-08-31 • Computation and Language
Computation and Language
AI summaryⓘ
The authors created a new test called READI to see how well computers understand indirect speech, which is when people say things that mean more than their words alone. Unlike past tests, READI combines pictures and conversations to help the computer figure out what is really meant based on the situation, especially in languages like Korean that rely heavily on context. They found that even the best current computer models have trouble understanding these indirect meanings when they rely on both images and dialogue. This shows there is still a need for better ways to teach computers how to read between the lines in different contexts.
Indirect Speech ActsPragmatic ReasoningMultimodal ModelsVisual ContextDialogue UnderstandingHigh-context LanguagesKorean LanguageVision-based Pragmatic Question AnsweringCross-lingual EvaluationNatural Language Processing
Authors
Jaehee Kim, Ji Hoon Chung, Seoyoon Park, Unsol Kim, Kyungwon Park, Ji Hak Kim, Yi-Jun Chen, Hansaem Kim
Abstract
Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.