Multi teacher distillation reveals factors shaping reinforcement learning updates
From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
Machine Learning
Summary
Combining the guidance from multiple AI teachers into one student model can be tricky, especially when these teachers learn by trial and error. The authors studied how different teacher signals change the student’s learning by looking closely at how the model’s parameters update in practice. They found that the way losses are averaged, the behavior of the optimizer, and number formats all affect how the student learns from its teachers. These insights help explain why some approaches to blending teacher advice work better than others in specific tasks like math.
What this means in practice
- •For machine learning engineers: Optimize student models by adjusting teacher signal combinations and averaging methods for better task accuracy in multi-teacher distillation setups.
- •For natural language processing developers: Improve multi-domain language model training by managing gradient behaviors and optimizer effects to retain strengths from multiple RL-trained teachers.
Authors
Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You
Abstract
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.