SparseOPD improves efficiency of on-policy distillation with selective corrections
Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
Computation and LanguageMachine Learning
Summary
When teaching AI models to generate better answers, a method called on-policy distillation uses corrections from a teacher model on all possible words. But checking every word takes a lot of memory for long answers. The researchers introduced SparseOPD, a smarter way that first looks at all corrections to pick only the important words before updating the model. This reduces memory use dramatically while keeping or improving the quality of the AI’s answers.
What this means in practice
- •For machine learning engineers: Reduce memory usage in training large language models with on-policy distillation while maintaining correction accuracy.
- •For ai model deployers: Efficiently update models in production that rely on teacher feedback to improve response quality without costly resource demands.
Authors
Yongliang Miao, Shuang Liu, Yanguang Liu, Yandong Bai, Mengnan Du
Abstract
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.