Visual token pruning improves efficiency in vision language models

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

Computer Vision and Pattern RecognitionMachine Learning

Summary

Large models that connect images and text use lots of computing power by looking at many parts of an image. The authors found that removing some parts too early based on text can make the model miss important details. They propose a method that first cuts down image parts using visual info, then later uses text to select the most relevant parts. This improves performance while saving computing resources on several tests.

What this means in practice

  • For machine learning engineers: Reduce computational cost in multi-modal AI systems by applying staged pruning to balance speed with accuracy for image-and-text tasks.
  • For mobile ai developers: Enable faster inference of vision-language models on resource-limited devices by pruning visual tokens with deferred textual guidance for better efficiency.

Authors

Minchan Kang, Kyeonghye Park, Seoyoung Cho, Daeshik Kim, Yucheol Cho

Abstract

Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT