KV-streams boost training speed for long-context agentic AI models
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
Machine LearningArtificial Intelligence
Summary
Large language models that act like agents need to remember long conversations or experiences, but this uses a lot of GPU memory and slows training down. The authors propose KV-streams, a way to keep memory use steady and speed training by streaming data forward instead of restarting it all the time. This makes training 2.6 to 5 times faster without hurting how well the model learns. They also found that this streamed memory helps the model remember important information for longer, improving its decision-making.
What this means in practice
- •For machine learning engineers: Speed up training of large agentic language models by integrating KV-streams to efficiently manage context memory without performance loss.
- •For software infrastructure teams: Reduce GPU resource demands for AI workloads involving long input histories by using KV-streams to optimize memory usage.
Authors
Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre Côté, Laurent Charlin, Guillaume Lajoie
Abstract
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.