Language model agents learn to use file memory for longer tasks
Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
Machine LearningArtificial IntelligenceComputation and Language
Summary
Long tasks can be hard for language AI agents because their memory is limited to what fits in a short context window. The authors created a new training setup called Coding Agent Memory Gym (CAMG), where agents can read, write, and search files as a kind of memory during their tasks. They trained these agents using reinforcement learning to use this file-based memory effectively across different types of tasks. Their best models perform competitively with much larger ones, showing that learning to manage memory using files helps agents handle longer tasks.
What this means in practice
- •For software developers: Build AI assistants that manage code and files over long development sessions, helping with remembering and revising work.
- •For automation engineers: Create automated systems that track and update persistent files to remember context in complex tasks beyond single interactions.
Authors
Lirui Luo, Kelong Mao, Heming Xia, Rongqing Li, Xinwei Yang, Luyu Chen, Kieran Wong, Yudong Guo, Xinrui Wang, Jiayin Zhu, Simiu Gu, Sulong Xu, Cong Fang
Abstract
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.