Papers for

multimodal ai product teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Vision language model rotation quantization improves accuracy at low bits

SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

Abstract: Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios and the positions of visual information, limiting statistical stability. Moreover, overly coarse aggregation through absolute values and averaging discards gradient signs and channel-wise differences, limiting the separation of modality-specific sensitivities. In contrast, the channel space provides a shared coordinate system across samples, making it a more natural basis for capturing stable task-sensitive structures. We therefore propose SubRot, a signed gradient subspace calibration method for VLM rotation quantization. Through eigendecomposition of the empirical Fisher matrix of activation gradients, SubRot identifies a sensitive channel subspace with three properties: cross-sample stability, clear sensitivity separation, and consistent signed effects on the autoregressive loss along certain directions. Guided by a local Taylor expansion, SubRot combines signed first-order guidance along sign-stable directions with second-order constraints along the remaining sensitive directions, while retaining MSE for overall reconstruction quality. This objective steers quantization errors toward loss-decreasing directions while controlling their magnitude. Experiments on five VLMs across five benchmarks show consistent average-score improvements over FlatQuant under W4A6 and W4A4, reaching 1.4 percentage points on LLaVA-NeXT-7B. Under W4A4, average accuracy degradation from FP16 remains within 1.4 percentage points across all evaluated models, while LLaVA-v1.5-13B exceeds its FP16 average score by 0.4 percentage points.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Reducing the size of vision-language models to deploy them efficiently can make them less accurate, especially when working with very small data sizes. The authors show this happens because current methods lose important directional information in the model’s gradients. Their method, SubRot, keeps track of these directions by finding stable patterns across many samples. This lets the model preserve accuracy better even when greatly compressed, as shown on several tests with different models.
Open → 2609.34884v1