Reward aligned weighting improves student model training accuracy

Reward-Aligned Reweighting for On-Policy Distillation

Machine Learning

Summary

Teaching a smaller language model to think like a bigger, smarter one usually treats all parts of an answer as equally important. The authors found this is not always the best way because some steps are more important to get right for finishing a good answer. They created a way to focus the teaching on the parts that really matter, based on how well the bigger model and smaller model agree on the final result. This method helped smaller models do better on math problems and coding tasks.

What this means in practice

  • For machine learning engineers: Improve training of smaller language models to better match larger models on tasks like math reasoning and code generation by focusing corrections where they matter most.
  • For software development tool builders: Enhance code generation quality in AI-powered programming assistants by using reward-aligned supervision to train compact models closer to expert outputs.

Authors

Haofeng Xu, Junwei Su, Lansong Diao, Wenchao Zhou, Chuan Wu

Abstract

On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R$^{2}$-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R$^{2}$-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.