Continuous context management reduces language model context size with tradeoffs
Continuous Context Management
Artificial Intelligence
Summary
Language model agents usually keep their full conversation history, which makes the context large and slow over time. The authors study a method called Continuous Context Management (CCM) that shrinks this history after every step, keeping only a summary and the latest observation. This approach cuts down how much information the model processes but can also reduce task success depending on the model used. They improve CCM by training models with extra feedback so CCM performs better, making it a promising way to keep language models efficient with less memory.
What this means in practice
- •For ai platform engineers: Deploy language model agents that maintain shorter interaction context to reduce prompt size and computational cost without full transcript retention.
- •For chatbot development teams: Implement continuous memory updates with reinforcement learning to improve chatbot efficiency while preserving task performance under limited context.
Authors
William Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan
Abstract
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.