Vision language models run faster by combining token pruning and quantization
P4Q: Co-designing Token Pruning and Quantization for Vision-Language Model Acceleration
Computer Vision and Pattern Recognition
Summary
Vision language models are powerful but usually require a lot of computing power and memory, making them hard to run efficiently. The authors show that combining two techniques—removing less important visual tokens and lowering the precision of numbers used during calculations—works better when done together rather than separately. They introduce P4Q, a method that coordinates both techniques to speed up the model while keeping its accuracy high. This approach can make vision language models almost three times faster without losing quality.
What this means in practice
- •For mobile app developers: Deploy vision language models efficiently on resource-limited mobile devices by jointly pruning and quantizing to reduce runtime without accuracy loss.
- •For cloud service providers: Accelerate multimodal model inference to reduce server costs and latency by co-designing token pruning with quantization calibration.$Commercial implications: Enables hosting faster and cheaper vision-language AI services with sustained accuracy, improving competitive cloud offerings.
Authors
Haizhao Jing, Zhenhao Shang, Haokui Zhang, Rong Xiao, Peng Wang
Abstract
Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence length and numerical precision. Existing workflows typically optimize these techniques independently or apply them sequentially. Their distinct optimization objectives leave critical interactions unaddressed and constrain the achievable compression performance. We revisit these designs and present P4Q, a practical co-design framework that jointly optimizes visual token pruning and low-bit quantization for efficient VLM inference. First, P4Q introduces a quantization-aware visual token selection strategy before the LLM. It applies fake quantization to copies of the features produced by the projector and selects visual tokens using statistics computed from these fake-quantized features, thereby conditioning the selector's feature-based decisions on simulated low-bit perturbations. Second, P4Q introduces a pruning-aware quantization calibration strategy. It uses the same selection strategy as pruning to calibrate the quantized model on the retained-token distribution, thereby aligning the calibration process with the pruned execution path used during deployment. By coupling these two components, P4Q achieves substantial inference speedups while maintaining comparable task performance, resulting in a better efficiency-accuracy trade-off than independently optimized pipelines. For instance, on LLaVA-NeXT, P4Q achieves an average end-to-end inference speedup of 2.8x across eight distinct test sets, while retaining higher accuracy than prior compression and quantization methods.