Models reveal how object roles and positions shape visual reasoning
Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning
Computation and LanguageComputer Vision and Pattern Recognition
Summary
Figuring out how objects relate to each other in space is tricky, especially when they swap places or switch roles in a question. The authors show how certain AI models track where objects are and understand their roles step by step. By tweaking parts of the models, they linked changes in remembering locations to swapping answers. They also found a way to adjust the models to make their reasoning about object positions more reliable, even when tested on new images.
What this means in practice
- •For computer vision developers: Improve AI models’ reasoning reliability about object positions in images by applying targeted adjustments to internal representations.
- •For natural language processing engineers: Use insights on role and position tracking to enhance language-based queries about spatial relationships in multimodal systems.
Authors
Yingjin Song, Denis Paperno, Albert Gatt
Abstract
High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models' behavior.