Large language models improve themselves by learning from their own reasoning
SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
Artificial IntelligenceComputation and Language
Summary
Improving large language models usually requires outside help like expert annotations or feedback, which is expensive and slow. The authors found that a single language model can use two ways of thinking: one that thinks deeply to generate detailed reasoning steps and another that gives quick answers. Their method teaches the model to learn from its own detailed reasoning steps to make better quick answers without needing external information. This approach helps the model get better on its own over time.
What this means in practice
- •For ai software engineers: Create language models that improve their reasoning skills during deployment without needing manual annotations or external feedback.
- •For natural language processing teams: Develop more efficient language model deployment by internalizing reasoning processes to boost performance without additional costly data.
Authors
Xiaoshu Chen, Xiangyu Wong, Sihang Zhou, Ke Liang, Xinwang Liu
Abstract
Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.