Vision language models improve reasoning with internal answer detection
From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models
Computer Vision and Pattern RecognitionMachine Learning
Summary
Vision-language models can answer simple questions about images but have a hard time when questions need multiple visual judgments combined. The authors studied how these models think internally by testing tasks like counting, recognizing shapes, and understanding spatial relations. They found that the models can form intermediate visual judgments without explicit reasoning and that the final answer becomes clear inside the model even before it finishes 'thinking.' They trained a small detector that predicts when the model is ready to answer, which reduced the amount of reasoning needed and improved accuracy on tested benchmarks.
What this means in practice
- •For ai system engineers: Reduce reasoning computation in vision-language models by detecting early answer readiness to improve efficiency and accuracy in visual question answering tasks.
- •For computer vision developers: Design visual reasoning components that combine judgments for complex multi-part questions by leveraging internal model states for better integration.
Authors
Rong Yu Xu, Prayag Tiwari, Shaolei Zhang
Abstract
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.