Vision language models improve reasoning by refining tokens differently

Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Some AI models that combine vision and language try to think better by changing the way their internal thoughts work during testing. Past methods treated all parts of these internal thoughts the same, which misses how some parts relate more to what is seen in images, while others involve harder reasoning steps. The authors introduced a method that looks at each part of the model's thought process differently and updates them based on whether they should focus more on visual clues or uncertain reasoning. This approach improved the accuracy of two popular vision-language AI models on several tests.

What this means in practice

  • For software engineers: Improve the inference accuracy of vision-language AI by selectively refining internal tokens based on their roles.
  • For ai developers: Build more accurate multimodal models that better combine visual information with reasoning during test-time.

Authors

Hao-Xuan Ma, Yihao Liu, Yutao Sun, Yanting Miao, Mengyu Zhou, YiCheng Xiao, Long Chen, Zhenguo Li, Han-Jia Ye, Xiaoxi Jiang, Guanjun Jiang

Abstract

Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.