Adaptive method improves large language model skill revision choices

StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills

Artificial Intelligence

Summary

Large language models can learn new skills by revising how they perform tasks without changing their internal settings. The challenge is deciding the best way to make these revisions because no single method is best for all tasks. The authors created StraTune, a system that picks the best revision method at each step by learning from previous attempts and results. This adaptive approach helped the models improve skills better than using fixed revision methods across a set of tests.

What this means in practice

  • For llm developers: Improve skill tuning strategies by adaptively selecting revision operators for better model performance on diverse tasks.
  • For nlp system engineers: Enhance task-specific language model adaptations without retraining by integrating an adaptive revision operator choice mechanism.

Authors

Zeping Liu, Yan Li, Ni Lao, Gil Wolff, Gengchen Mai

Abstract

Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision forms applied under it. However, we observe that no single revision operator consistently performs best across tasks, and repeatedly applying an unsuitable operator can limit further improvement. We propose StraTune (strategy-guided skill tuning), which lets a frozen optimizer LLM choose the revision operator at every round from the optimization state, which is defined as the current execution feedback together with the recorded outcomes of earlier strategies and forms. Candidate skills from every revision operator pass one candidate evaluation, which screens for gains and regressions on a small sample set and validates them on a larger one, and every outcome is written back to the optimization state for later choices. Across four benchmarks and two LLM settings, StraTune outperforms all five baselines in most settings. Ablations attribute the gains to the adaptive choice of the revision operator, since fixed, random, scheduled, and bandit strategy choices all score lower, and skills learned with a small target LLM also improve a stronger one. Code and learned skills are available at https://github.com/seai-lab/StraTune.