Papers for

ai serving engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

StepQuant reduces memory use for AI with smarter state compression

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Abstract: Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.

Tue 29 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Modern AI language models remember information during output generation using states that can take up a lot of memory. The authors found that errors in compressing these states affect the model’s accuracy differently depending on when and where they happen. They created STEPQuant, a method that decides how much detail to keep in each part of these states to reduce memory without losing accuracy. Tests showed STEPQuant keeps the same quality as the original but uses much less memory, making AI models easier to run efficiently.
Open → 2609.38169v1