FlexComp adapts context compression ratio per input for better efficiency

FlexComp: One Model for Every Ratio in Context Compression

Computation and Language

Summary

Large language models often need to compress long pieces of text into smaller chunks they can process. Traditionally, one setting for how much to compress is fixed for all inputs, which isn’t efficient because some texts need more detail than others. The authors of this paper developed FlexComp, a single model that can compress texts at varying levels of detail depending on the text itself. This means it saves memory and speeds up processing without losing much accuracy, adapting dynamically for each input.

What this means in practice

  • For machine learning engineers: Optimize memory use and decoding speed when deploying large language models by dynamically adjusting context compression per input.
  • For cloud service providers: Reduce resource consumption in LLM serving infrastructures by automatically selecting compression levels to balance accuracy and efficiency.

Authors

Kaiyan Zhao, Zhongtao Miao, Akiko Aizawa, Yoshimasa Tsuruoka

Abstract

Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chosen ratio is applied uniformly to all inputs, whose actual needs vary drastically. We propose FlexComp, a method-agnostic framework that decouples the ratio from both training and deployment: Matryoshka-style training samples the memory budget $K$ per instance, turning one model into an any-ratio compressor, and the budget is then chosen per input by: (1) confidence-based cascade routing or (2) a lightweight learned $K$ predictor. Across ICAE, 500xCompressor, and SAC on MRQA, a single FlexComp model matches separately trained fixed-ratio specialists with minimal degradation. Cascade routing preserves over 98% of the mildest ratio's accuracy at up to 266x average compression; the $K$ predictor, in a single compression-decoding pass, reaches 158-236x within 0.7 F1 of the mildest ratio. At serving-scale batch sizes, the $K$ predictor cuts context KV cache by 50% and improves decoding throughput by 47%.