Vision transformers keep accuracy when compressed with global local features

GLF-Q: Global-Local Feature-based Quantization for Vision Transformers

Computer Vision and Pattern Recognition

Summary

Vision transformers are powerful image recognition models but shrink them too much, and their accuracy drops sharply. The authors introduce GLF-Q, a technique that helps compress these models to very small sizes without losing much accuracy by aligning detailed and overall features in the model’s layers. They also apply a special transform to spread out odd activation values, making the compression simpler and more effective. Their method works well across various models and keeps performance reliable even with less common input data.

What this means in practice

  • For machine learning engineers: Compress vision transformer models to 3-bit precision while maintaining high image classification accuracy for deployment on resource-limited devices.
  • For gpu software developers: Implement 8-bit quantized vision transformer inference optimized for faster execution on GPUs with negligible accuracy loss.

Authors

Peilin Sun, Guang Liang, Jin Tong, Jianxin Wu

Abstract

Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervision is prone to overfitting on limited calibration data, while Hessian approximations incur structural truncation errors. To address these limitations, we propose \textbf{GLF-Q}, a novel PTQ framework guided by Global-Local Feature alignment. GLF-Q propagates quantized block outputs through downstream full-precision layers to align penultimate-layer representations under local output regularization, providing downstream feature supervision without explicitly approximating the Hessian or using a Taylor expansion. Furthermore, offline Hadamard transformations are introduced with zero runtime overhead to disperse activation outliers across channels, effectively contracting dynamic ranges and reducing quantization errors. Meanwhile, optimizing this loss via a Straight-Through Estimator (STE) achieves rapid convergence, bypassing continuous relaxation rounding formulations such as AdaRound. Extensive experiments across representative ViT architectures demonstrate that GLF-Q with standard uniform quantizers substantially outperforms state-of-the-art methods under 3-bit quantization on image classification. In addition, GLF-Q exhibits strong out-of-domain calibration robustness and achieves speedups under 8-bit GPU deployment.