Kv cache quantization reduces memory for long context language models
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Machine Learning
Summary
Long conversations with AI need a lot of memory to remember what was said before, which slows things down. The authors came up with a way called WUSH-KV that shrinks this memory by using smart math based on the data’s patterns, so less space is needed without losing much information. They tested it on different language models, and it worked as well or better than other methods even when squeezing the memory to just 2 bits of data. This makes it easier to run AI models on long texts efficiently.
What this means in practice
- •For natural language processing engineers: Reduce memory footprint and bandwidth needs when running long-context transformer models in production by applying WUSH-KV quantization to KV cache.
- •For machine learning infrastructure teams: Integrate data-adaptive low-bit quantization transforms into inference pipelines to improve efficiency and reduce hardware usage for large language model deployments.
Authors
Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler, Dan Alistarh
Abstract
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.