Kv cache compression improves large language model efficiency
D-Quant: Driftable Entropy Coding for KV Cache Quantization
Computation and Language
Summary
Large language models use a key-value (KV) cache to speed up processing, but this cache takes up a lot of memory, which slows things down. Most methods to reduce this cache's size use fixed-bit quantization, which doesn't work well because it treats frequent and rare data equally and loses important information. The authors developed D-Quant, a new way to compress the KV cache that uses entropy coding to save memory without slowing down processing, by converting variable-length codes into fixed-size pieces that computers can handle efficiently.
What this means in practice
- •For machine learning engineers: Reduce memory usage of KV caches in large language models to improve inference efficiency without sacrificing accuracy.
- •For hardware accelerator developers: Design memory access patterns in attention kernels to handle compressed KV caches with fixed-size bitstreams for faster parallel processing.
Authors
Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu
Abstract
The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is particularly attractive due to its effectiveness and ease of deployment. However, most existing methods rely on fixed-width quantization, where a $b$ bit representation is inherently limited to $2^b$ quantization levels. As the bit width decreases, the number of available levels shrinks exponentially, leading to severe information loss and rapid performance degradation. We further observe that fixed-width quantization fails to exploit the highly non-uniform distribution of KV cache. After rotation and normalization, KV values approximately follow a normal distribution, with most values concentrated near the center and only a small fraction appearing in the tails. Nevertheless, fixed-width coding allocates the same number of bits to frequent and rare symbols. Entropy coding naturally exploits such non-uniformity by assigning shorter codewords to frequent symbols and longer ones to rare symbols, substantially reducing the average number of bits required for representation. However, its variable-length output is not suited to highly parallel attention kernels, where efficient dequantization and computation rely on regular memory layouts and fixed-stride accesses. To bridge this gap, we propose \textbf{D-Quant}, a flexible KV cache quantization framework that introduces a \textbf{drift} mechanism to convert entropy-coded representations of each token into fixed-size bitstreams, enabling regular memory access and parallel dequantization within attention kernels.