Looped transformers quantization fails without feedback and calibration fixes
Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
Machine LearningComputation and Language
Summary
Quantizing looped transformers—in which the same weights are reused repeatedly—can cause errors that get worse over time. The authors found two main problems: one where errors build up because a key part lacks a direct 'identity path' to keep things stable, and another where the calibration method misses important later signals. They showed that by improving how calibration collects information across multiple steps, this error can be greatly reduced, matching higher-precision results. These insights help make very low-bit quantization of looped models more reliable.
What this means in practice
- •For ai hardware engineers: Improve low-bit quantization of looped transformer models by addressing feedback and calibration blind spots to reduce accuracy loss in efficient AI accelerators.
- •For machine learning engineers: Optimize post-training quantization workflows for looped transformer architectures to maintain model accuracy without costly retraining.
Authors
Nux Li
Abstract
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Controlled experiments on linear filters and Mamba state-space models show that feedback exposure also occurs outside transformers. Grouped INT4 reveals a separate failure, calibration blindness: our one-step GPTQ baseline builds its Hessian from step-0 activations, leaving input directions used later in the recurrence nearly unweighted. Across nine checkpoints from seven looped architectures, one-step GPTQ is worse than round-to-nearest (RTN) on the primary task metric for five checkpoints. Accumulating the GPTQ Hessian across recurrence steps outperforms both one-step GPTQ and RTN on all nine checkpoints and recovers bf16-level accuracy on Huginn. These results separate two questions for PTQ on looped models: where quantization error enters the recurrence, and which states calibration sees.