Architecture enables AI agents to work on tasks lasting many days
An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
Artificial IntelligenceMachine Learning
Summary
Some AI helpers need to do jobs that take many days or weeks, but they forget too much and can’t keep track over time. The authors show that these helpers need a special system around the AI to remember and keep learning without losing old knowledge. They made a setup with three parts: different time levels that summarize what happened before, regular actions called ticks, and a way to pass problems up to smarter parts only if needed. They tested their system on a ten-day task where a human checked in just once a day, and it kept track and improved without changing the core AI.
What this means in practice
- •For automation engineers: Run AI systems that continuously manage complex processes over days with minimal human oversight using hierarchical time levels and periodic action ticks.
- •For software operations teams: Build AI agents that handle long-term remediation tasks by escalating issues through successively capable models only when needed.
Authors
Erik Nijkamp, Anurag Koul, Egor Pakhomov, Bo Pang
Abstract
Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually. This ability lies in the harness around the model rather than in the model itself. We derive seven bottlenecks from the long-horizon setting and answer them with a hierarchical architecture of three parts: (i) levels indexed by time scale, each keeping a bounded file summarising the level below; (ii) a clocked tick as the unit of autonomous action; and (iii) cascaded intelligence, where work is escalated to a more capable model only after failing review. We report on a ten-day campaign in which an agent built on this architecture reproduced a published reinforcement-learning result with a human attending once a day, and show (1) the agent kept the thread across every context reset and session boundary of the campaign, (2) operating knowledge written early changed later behaviour with no change to model weights, and (3) where learned components would enter such a system. Overall, our experience suggests continual learning for these agents needs a substrate outliving every context and process, and the checks the harness already runs are where a learner belongs.