Self evolving agents learn precise control updates inside workflows

The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents

Computation and Language

Summary

Self-evolving agents usually learn from past actions, but struggle to apply lessons exactly where they help most in a multi-step task. The authors propose EvoCUE, a system that represents an agent’s workflow as a state machine and learns to make precise control changes at specific steps based on past successes or mistakes. By testing edits right where they apply, EvoCUE improves the agent’s future task completion without causing distractions elsewhere. This approach helps agents better transfer knowledge in complex, multi-tool tasks.

What this means in practice

  • For enterprise automation teams: Customize complex office workflow automation by learning precise control adjustments to meet organizational rules across similar tasks.
  • For software developers: Improve multi-step AI tool integration by enabling agents to learn specific control updates at exact workflow steps from past executions.

Authors

Yunhe Su, ZiYi Dong, Tong Yu, Weijian Deng, Hao Li, Bowen Jiang, Pengxu Wei

Abstract

Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experience reuse a problem of localized control rather than memory alone. We introduce EvoCUE (Evolution through Control Updates from Evidence), a framework for learning reusable control-program updates from completed agent executions. EvoCUE represents the agent as an explicit state-machine controller, whose nodes perform model or tool calls and whose edges define where control passes next. This makes the workflow editable at precise locations, so each learned update can specify what to add, where it acts, and when it applies. From completed trajectories, EvoCUE uses residual goals and observed execution traces to propose localized instruction or skill edits. Each candidate is evaluated at the point where it would act by resuming the parent and edited controllers from the same checkpoint and comparing their final outcomes. Accepted edits are compiled with applicability rules, confirmed on held-out tasks, and inherited by later executions. We evaluate EvoCUE on long tool-use environments where learned conventions must reach the right execution step. From a minimal AppWorld controller without benchmark-specific onboarding instructions, EvoCUE learns the missing task-completion convention and substantially improves success on Test-Normal and Test-Challenge. On PAST-Bench office workflows, EvoCUE transfers organizational requirements from prior episodes to later tasks, improving task-execution quality. These results show that self-evolving agents should place experience inside the control flow, rather than only store it as text.