Large language models improve skills by fixing mistakes locally
A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents
Artificial Intelligence
Summary
Complex tasks can often be done in many right ways, so forcing a language model to copy one fixed correct solution doesn’t help it learn well. The authors found that even when a model’s attempt fails, it often makes good progress before going off track. They developed SkillPivot, which identifies where the model’s reasoning starts to fail and then uses a stronger model to fix just that part. This focused approach helps the model improve without losing what it already does correctly.
What this means in practice
- •For ai application developers: Improve language model agents' problem-solving skills on complex, multi-step tool use tasks by refining only the failed parts of their reasoning paths.
- •For automation engineers: Boost performance of autonomous agents performing multi-step operations by targeted self-correction where errors start, preserving earlier effective actions.
Authors
Yichun Feng, Jiawei Wang, Haozhe Sun
Abstract
Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify where productive problem solving begins to break down, rather than reflect coarsely over the entire failure. Based on this insight, we propose SkillPivot, a deviation-point-guided framework for skill self-evolution. SkillPivot detects the transition from a useful prefix to an erroneous suffix using execution validity, goal progress, and action diversity. A stronger teacher then continues from the same prefix and produces a successful alternative under the same interaction history. By contrasting the student's failed suffix with the teacher's successful suffix, SkillPivot generates localized skill updates while preserving already effective guidance. Experiments on ToolQA, LogicBench, and WildClawBench show that SkillPivot consistently outperforms competing skill-evolution methods, improves multiple agent models, and produces compact, transferable skill updates.