Self evolving language models learn faster without challenger training loop

Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

Machine LearningArtificial Intelligence

Summary

Language models can train themselves by creating tasks and learning from their own answers, but usually this needs a separate training step for the task generator. The authors propose a new method called Direct Self-Evolving Optimization (DEO) that removes this extra training step by letting the solver guide task selection directly. Their method trains only the solver model while using a fixed task generator and still improves reasoning skills efficiently. Experiments show DEO matches previous methods while using less training time and even benefits from using a frozen external language model to generate tasks.

What this means in practice

Authors

Yuyang Deng, Yu Wang, Jiayun Wang

Abstract

Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over $50\%$ less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.