Method helps stronger AI models learn better from weaker ones
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Machine LearningComputation and Language
Summary
Sometimes, weaker AI models teach stronger models, but the stronger ones don’t always improve beyond the weaker teachers. The authors present a method called On-Policy Reverse Distillation (OPRD) that helps stronger models learn faster and perform better by focusing on directions where they can improve beyond their teachers. Instead of copying weak models exactly, OPRD uses feedback from how the student model performs to guide learning, allowing it to surpass the weak teacher’s abilities. This approach works well when teaching models in a sequence or from multiple weak teachers and also helps when learning goes from strong to weak models.
generalizationdistillationpolicy optimizationteacher modelstudent modelreinforcement learningpolicy gradientmodel capacitymulti-domain learning
Authors
Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun
Abstract
Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.