Moscopt improves llm agents by mixing and choosing skills dynamically

MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents

Artificial IntelligenceComputation and Language

Summary

LLM agents use language prompts called skills to perform tasks, but existing approaches optimize only one skill at a time, missing benefits from combining multiple strategies. The authors propose MOSCOPT, a method that optimizes a group of skills together and learns when to activate which skills dynamically. They introduce a new optimization technique called EditAdam to improve these skills without needing parameter tuning. Experiments show MOSCOPT performs better than previous methods by using a mix of skills and updating them collectively in an organized way.

What this means in practice

  • For llm system developers: Improve performance of multi-skill language model agents by jointly optimizing skill pools with dynamic selection strategies.
  • For conversational ai teams: Deploy adaptive language bots that select and combine multiple conversational strategies optimally for better user interaction.

Authors

Zhenyu Zhang1, Jiudong Yang

Abstract

Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template---missing the synergy among multiple complementary strategies. We propose MOSCOPT, a text-native, parameter-free algorithm that jointly optimizes a pool of $N$ skills and a gating skill $G$ that dynamically selects $K$ skills per step. To effectively optimize the skills, we build the EditAdam with internally maintained dual states. Through the three-phase interleaved updates with EditAdam, the system monotonically improves without gradient or parameter tuning. Extensive experiments and detailed ablations across 5 benchmarks and 3 target LLMs demonstrate that MOSCOPT consistently outperforms all baselines, and confirm that both the mixture-of-skills architecture with selective activation and the collective evolution with three-phase interleaving are essential to its superior performance. Code is released https://github.com/zhangzhenyu13/SummerClaw/tree/master/summerclaw/agent_trainer/algorithms/moscopt.