Scalable memory method boosts AI learning for long tasks

Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation

Machine Learning

Summary

In many AI tasks, it's hard to learn from a lot of past experience because remembering everything is slow and uses a lot of memory. The authors introduce a new method called Recurrent Algorithm Distillation that helps AI keep track of important past information by squeezing it into a smaller summary, so it can learn efficiently from a long history without needing lots of memory. Their system combines this compressed memory with recent new information to decide what to do next. Experiments show this approach works just as well as older methods but needs less memory, making it easier to use on more complex problems.

What this means in practice

  • For robotics engineers: Improve robot decision-making in long and complex tasks by efficiently summarizing past experiences into manageable memory representations.
  • For game ai developers: Enable game AI agents to remember and act on long sequences of past events without large memory requirements, enhancing long-term strategy learning.

Authors

Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao

Abstract

Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.