SkillFocus improves evolving AI agents by separating task needs from behaviors
SkillFocus: Evolving Agent Skills via Capability Decomposition
Artificial Intelligence
Summary
Improving AI agents that perform tasks repeatedly can be tricky because their updates depend on how they acted before. The authors propose SkillFocus, which breaks down what tasks need into fixed capabilities, keeping these needs clear even as the AI skill changes. This method helps identify which capability causes most problems and focuses updates there using the right evidence. The authors show that SkillFocus works better and more efficiently than older methods on several task tests.
What this means in practice
- •For software engineers: Enhance AI agents in complex software to update skills more effectively by focusing revision on core task capabilities causing failures.
- •For automation strategists: Improve automation workflows by evolving intelligent agents that better generalize across varied tasks using fixed capability structures.
Authors
Ning Wang, Zhiren Gong, Bingdong Li, Peng Yang, Aimin Zhou
Abstract
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24\% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task--capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.