Coding agent learns when and how to compact context for better long tasks
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
Computation and Language
Summary
Coding agents work on software tasks that need many steps, but remembering too much can be a problem as they go. The authors created AutoCompact, which helps these agents decide when to shrink their memory and what details to keep for later. They trained the agent by correcting poor memory choices and then improved it with rewards when the overall task succeeded. This method led to better performance in coding tests, even when the agent's memory was limited or very large.
What this means in practice
- •For software development teams: Improve automated coding tools by enabling agents to manage memory efficiently during complex long coding tasks.
- •For automated code testing teams: Enhance long-sequence code inspection and testing agents to maintain relevant context without overflow or loss.
Authors
Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong
Abstract
Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.