Context evolution improves long task solving in AI agents
Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
Artificial Intelligence
Summary
Large language model (LLM) agents often struggle with long tasks because the information they rely on grows too big and messy over time. The authors developed ContextEvo, a system that learns how to manage what information the AI focuses on by reviewing past task attempts and fixing mistakes related to context. This helps the AI keep only the most useful information when making decisions. Tests showed that ContextEvo helps AI agents do better on long, complex tasks than some earlier systems.
What this means in practice
- •For ai platform engineers: Improve AI agent frameworks by integrating adaptive context management policies to enhance performance on lengthy, multi-step tasks.
- •For software tool developers: Develop smarter automation tools that manage accumulated information effectively, enabling more reliable completion of complex workflows.
Authors
Weiyuan Li, Jinghan Xu, Aili Chen, Xintao Wang, Shuang Liang, Jiaqing Liang, Deqing Yang
Abstract
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce ContextEvo, a framework that learns a context policy from long-horizon trajectories. ContextEvo reconstructs the model-visible context at key decision points, identifies context-related failures, and applies targeted policy updates. Starting from the open-source Pi-agent harness, ContextEvo improves performance across three long-horizon task benchmarks, achieving results comparable to or better than several prominent agent harnesses, including Codex, OpenCode, and OpenClaw. Additional analyses show that fixed or locally evolved context strategies can fall short under long-horizon information pressure, while our methods adapt to the information demands of each environment.