Policy distillation improves reinforcement learning outcomes beyond initial accuracy
RL Starts before RL: On Policy Distillation for Better Reinforcement Learning
Machine Learning
Summary
Reinforcement learning (RL) helps machines learn by trial and error, but it matters which starting point you choose. The authors studied a way to prepare for RL called on-policy distillation (OPD), where a student learns from a teacher’s behaviors before actual RL training. They found that starting with OPD can lead to better final performance, even if it doesn’t improve the student’s accuracy right away. This improvement might come from the student better matching the teacher’s range of good behaviors, giving RL more options to improve upon.
What this means in practice
- •For machine learning engineers: Improve reinforcement learning model training by initializing with on-policy distillation to achieve higher final performance.
- •For ai application developers: Create more reliable reinforcement learning systems by selecting distillation methods based on expected downstream training performance.
Authors
Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Tian, Sirry Chen, Xingyu Liu, Xiangnan Wu, Jiawei Guo, Haowen Hou, LingHan Chen, Zhongyu Wei, Jiaqi Wang
Abstract
Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students initialized with OPD reach higher final performance than those trained with direct RL or supervised fine-tuning followed by RL. This advantage can emerge even when OPD produces little immediate improvement in accuracy. Pre-RL Pass@k does not fully explain the benefit: similar or even higher values do not necessarily lead to better performance after RL. Behavioral analyses point to alignment with the teacher's distribution beyond top-1 agreement as a possible explanation. Such alignment may favor higher-quality reasoning paths while retaining alternatives that RL can further refine using outcome feedback. We further examine how trajectory sources and divergence objectives affect the value of distillation for subsequent RL. Standard reverse-KL OPD performs better before RL, but forward-KL OPD overtakes it afterward; with teacher-generated distillation trajectories, reverse KL remains ahead at both stages. These findings suggest that the preferred distillation objective depends on both the trajectory source and the training that follows. Our results support evaluating OPD as preparation for RL and selecting distillation choices by the performance achieved after subsequent training.