Adaptive pruning cuts visual tokens to speed up multimodal language models
StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models
Computer Vision and Pattern Recognition
Summary
Multimodal language models combine text and images but spend a lot of time processing visual tokens. The authors propose StepPrune, a method that smartly selects which visual tokens to keep based on what was selected before and the text context. This method decides how many tokens to keep dynamically instead of a fixed amount and speeds up processing without much loss in performance. StepPrune also handles high-resolution images by grouping tokens for selection.
What this means in practice
- •For mobile app developers: Reduce visual token processing time in apps using multimodal language models for faster responses.
- •For chatbot platform engineers: Improve inference speed of multimodal chatbots by adaptively pruning visual tokens without severely impacting accuracy.
Authors
Hansen Zhang, Landi He, Mingde Yao, Lijian Xu
Abstract
Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selection-dependent interactions and variations in visual complexity across samples. In contrast, we propose StepPrune, which formulates visual-token pruning as an adaptive sequential decision process. Conditioned on previously selected tokens and textual context, StepPrune progressively constructs the retained subset and automatically determines its size through a learned STOP action. During training, a variance-preserving noise gate provides a differentiable surrogate for the discrete selection process, whereas during inference, unselected tokens are physically removed before language-model prefill. A grouped selection mechanism further extends StepPrune to high-resolution inputs. Experiments across LLaVA-1.5, LLaVA-NeXT, Qwen2.5-VL, and InternVL3 show that StepPrune achieves the best average normalized performance retention across all evaluated pruning rates on LLaVA-1.5, Qwen2.5-VL, and InternVL3, while remaining competitive on the substantially longer AnyRes prefixes of LLaVA-NeXT. On LLaVA-1.5, StepPrune retains 94.6% of the full-prefix normalized performance while pruning 88.9% of the visual tokens. At a mean retained count of 64, StepPrune reduces prefill latency from 59.95 ms to 40.05 ms, corresponding to a 1.50x prefill speed-up.