Balanced fitting improves post-training quantization for vision-language models

Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

Computer Vision and Pattern RecognitionMachine Learning

Summary

Making large vision-language AI models smaller and faster is important but tricky because they need to work well on many different tasks. The authors found that just trying to copy the full-size model exactly when shrinking it isn't always best. Instead, their new method called Balanced Fitting smartly balances copying important parts closely while letting other parts simplify more to help the model perform better overall. This approach works better than older methods on several different large models.

What this means in practice

  • For machine learning engineers: Deploy large vision-language models more efficiently on devices with limited computing power without losing performance on diverse tasks.
  • For ai software developers: Improve the accuracy of compressed vision-language AI in applications like image captioning and text-based image search.

Authors

Minchan Kang, Kyeonghye Park, Seungyeon Sa, Seoyoung Cho, Daeshik Kim, Yucheol Cho

Abstract

Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize reconstruction loss with respect to the full-precision model, potentially over-preserving FP behavior and calibration-specific bias. Rather than treating quantization solely as an error to be minimized, we observe that it can also provide beneficial regularization for certain layers and modalities. Motivated by this observation, we propose Balanced Fitting, a quantization effect-based framework that balances precision and regularization beyond reconstruction-based optimization. By measuring layer- and component-wise quantization effects for weights, vision activations, and text activations, Balanced Fitting combines fine-grained fitting for sensitive components with coarser fitting to exploit potential regularization benefits. Experiments on multiple LVLMs show that our method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance. The source code is publicly available at https://github.com/kmc3661/BFQ