Mismatch Matters: On-Policy Distillation Beyond Token Agreement

2026-08-10Artificial Intelligence

Artificial IntelligenceComputation and Language
AI summary

The authors studied a problem in on-policy distillation (OPD) where the student model tricks the system by copying tokens to match the teacher model but still gives bad answers. They identified two main token mismatches: extra tokens the student adds that the teacher doesn’t approve, and missing tokens the teacher prefers but the student doesn’t use. To fix this, they developed TIDE, which limits excessive tokens and adds back missing important tokens without needing the student to guess them first. Their method improved performance on reasoning tasks, especially when the student and teacher models were quite different.

On-policy distillationLanguage modelsTeacher-student mismatchToken-level correctionHellinger shapingTop-K samplingMathematical reasoning benchmarksModel alignmentQwen3 modelsReward shaping
Authors
Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou
Abstract
On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shift our focus from agreement to teacher-student mismatch, and find that mismatch tokens can be mainly categorized into two types: student-excess tokens and student-deficit tokens. Student-excess tokens are generated by the student but assigned near-zero probability by the teacher; their log-ratio corrections grow unbounded and destabilize the update. Student-deficit tokens, in contrast, are preferred by the teacher but rarely sampled by the student; their absence blocks the transfer of the teacher's reasoning patterns. To tackle these mismatch directions, we propose TIDE (Token-level Independent Deficit-Excess correction), which applies bounded Hellinger shaping to suppress the most severe sampled excesses and an analytic teacher top-$K$ injection to restore deficient probability mass without requiring deficit tokens to be sampled. Across mathematical reasoning benchmarks with multiple Qwen3 teacher-student pairs, TIDE consistently outperforms standard OPD and recent token-selection and reward-shaping baselines. Moreover, the gains of TIDE are more pronounced under strong teacher-student mismatch, where it improves Avg@8 from 6.9% to 20.3%, reduces average response length by a factor of 3.6, and substantially reduces formatting failures. Code is available at https://github.com/yzc-666/TIDE