Quantizing neural networks with improved error control using regularization
Generalization behavior of OPTQ and the role of regularization
Machine Learning
Summary
Big AI models can be made smaller and faster by turning their many numbers into fewer bits, but this can cause mistakes. The authors study a way called OPTQ that carefully picks how to round these numbers to keep errors low on example data. They prove math results that explain how well this method works on new, unseen data, showing an important role for a penalty term called regularization. Using these insights, they suggest a better value for this penalty and show it works well in tests compared to older suggestions.
What this means in practice
- •For machine learning engineers: Use the recommended regularization setting to quantize large models with better error control for deployment on limited hardware.
- •For edge device developers: Implement quantized neural networks that maintain accuracy when tested in real environments by applying improved regularization choices.
Authors
Erin George, Rayan Saab
Abstract
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive bounds for the expected squared error accrued by the algorithm when a test point is drawn from a fixed distribution. We prove two results. One result relates the generalization error to the error on a calibration dataset comprising independent samples from the same distribution as the test distribution. The other result bounds the generalization error of stochastic OPTQ for all sufficiently nice distributions, regardless of the calibration dataset. In both of these results, the regularization term $λ$ plays an important role. We use insights from these results to make a new recommendation for the choice of $λ$ and see that this choice of $λ$ preforms favorably in experiments when compared to prior recommendations in the literature.