QuantForge improves four-bit quantization for large language models
QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization
Machine Learning
Summary
Reducing the memory needed by large language models usually harms their accuracy, especially when using very tight four-bit quantization. The authors created QuantForge, a system that helps discover better quantization methods by carefully analyzing errors and revising algorithms step-by-step. Their system guided the discovery of HiRes, a new quantizer that better preserves accuracy by adjusting how errors and residues propagate through the model. Tests on multiple tasks show that HiRes achieves more accurate results compared to other quantizers at the same memory budget.
What this means in practice
- •For machine learning engineers: Implement more accurate four-bit quantization to reduce memory footprint of large language models while maintaining prediction quality.
- •For data center infrastructure teams: Optimize resource usage by deploying large language models with aggressive quantization that balances accuracy and memory demands.
Authors
Qiulin Shang, Zhoutong Wu, Jie Hu, Kun Yuan
Abstract
Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the network. The useful algorithmic decomposition is therefore not fully known before search. LLM-driven program evolution offers a way to explore these choices, but performance scores alone do not explain which design should change next. We introduce QuantForge, a PTQ discovery system that records competing explanations, selects controls that distinguish them, and checks that successor code implements the resulting conclusions. This residual compilation guides program revisions while retaining useful programs even when their original explanations are rejected. Remeasuring the revised program reveals the next error to address. This process discovers HiRes, a fixed MXFP4 quantizer that shapes coordinates, refines legal code assignments, and recovers errors along attention and MLP paths. Each stage acts on residuals measured after the preceding stage has executed. Across seven tasks, HiRes achieves the lowest seven-model Robust Fit (0.09300) and the lowest quantized Fit-7 at 32B. In matched-budget comparisons of LLM-driven program evolution, each with 240 evaluator calls, QuantForge reaches a held-out transfer target in six of eight runs, compared with three each for textual memory and reflection memory, and one for score-only evolution, despite evaluating fewer new programs. These results show that QuantForge improves the discovery of transferable PTQ algorithms by turning controlled evidence into subsequent program changes.