Recursive skill evolution improves large language model agents performance

R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution

Artificial Intelligence

Summary

Large language model (LLM) agents can get better at tasks by reusing and improving the skills they combine to solve problems. The authors present a method called R² Flow that helps these agents improve themselves step-by-step, by learning, verifying improvements, and updating their skill libraries more reliably. This method organizes how skills are combined so the system can share learning across similar task paths and decide more accurately which skills to change. In tests involving question answering, math reasoning, decision making, and code generation, R² Flow led to better accuracy and more precise skill updates than previous approaches.

What this means in practice

Authors

Mingda Zhang, Qiang Huang, Yanjin Li, Zijia Wang, Qika Lin, Xiaoying Tang, Tiesunlong Shen

Abstract

LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at https://github.com/beita6969/r2flow.