Vision language model rotation quantization improves accuracy at low bits
SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization
Computer Vision and Pattern Recognition
Summary
Reducing the size of vision-language models to deploy them efficiently can make them less accurate, especially when working with very small data sizes. The authors show this happens because current methods lose important directional information in the model’s gradients. Their method, SubRot, keeps track of these directions by finding stable patterns across many samples. This lets the model preserve accuracy better even when greatly compressed, as shown on several tests with different models.
What this means in practice
- •For machine learning engineers: Improve deployment efficiency of vision-language models on low-memory devices by preserving accuracy during extreme quantization.
- •For ai hardware developers: Design quantization techniques that direct errors toward loss-improving parameter directions to maintain performance in vision-language chips.
- •For multimodal ai product teams: Build lightweight vision-language applications that deliver near full-precision accuracy at significantly reduced compute cost.$Commercial implications: Enables commercial products that run large VLMs efficiently on edge devices with limited resources, creating new market opportunities.
Authors
Zhenhao Shang, Haizhao Jing, Haokui Zhang, Guoting Wei, Rong Xiao, Jianqing Gao, Peng Wang
Abstract
Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios and the positions of visual information, limiting statistical stability. Moreover, overly coarse aggregation through absolute values and averaging discards gradient signs and channel-wise differences, limiting the separation of modality-specific sensitivities. In contrast, the channel space provides a shared coordinate system across samples, making it a more natural basis for capturing stable task-sensitive structures. We therefore propose SubRot, a signed gradient subspace calibration method for VLM rotation quantization. Through eigendecomposition of the empirical Fisher matrix of activation gradients, SubRot identifies a sensitive channel subspace with three properties: cross-sample stability, clear sensitivity separation, and consistent signed effects on the autoregressive loss along certain directions. Guided by a local Taylor expansion, SubRot combines signed first-order guidance along sign-stable directions with second-order constraints along the remaining sensitive directions, while retaining MSE for overall reconstruction quality. This objective steers quantization errors toward loss-decreasing directions while controlling their magnitude. Experiments on five VLMs across five benchmarks show consistent average-score improvements over FlatQuant under W4A6 and W4A4, reaching 1.4 percentage points on LLaVA-NeXT-7B. Under W4A4, average accuracy degradation from FP16 remains within 1.4 percentage points across all evaluated models, while LLaVA-v1.5-13B exceeds its FP16 average score by 0.4 percentage points.