Cross tokenizer training improves language models with better event completion

Beyond Token Alignment: Event Completion for Cross-Tokenizer On-Policy Distillation

Computation and LanguageArtificial Intelligence

Summary

Language models often learn from each other during training, but when they use different ways of breaking text into pieces (tokenizers), this can cause problems. The authors studied how a student model can better learn the correct way to finish parts of words when the teacher model’s pieces don’t match exactly. They introduced a new method called Event-Set Completion Distillation that looks at all the possible correct ways to complete pieces and trains the student model accordingly. This approach helps the student better match the teacher’s knowledge without needing extra data or changes to the model’s vocabulary.

What this means in practice

  • For language model engineers: Improve training of language models that use different tokenizers by aligning event completion probabilities for better knowledge transfer.
  • For code generation developers: Enhance code generation models by refining how smaller models learn from large teachers when token splits differ, resulting in better completion accuracy.

Authors

Jiacheng Liu, Jingwei Song, Qituan Zhang, Siheng Chen, Linfeng Zhang

Abstract

On-policy distillation (OPD) transfers knowledge between language models through teacher supervision on student-generated trajectories. With different tokenizers, a single teacher token may require multiple student tokens to generate, creating intermediate states where the event is entered but not yet completed. Existing cross-tokenizer methods align tokens or text spans to construct comparable prediction targets. We study a complementary problem after partial generation: once the student produces a prefix of a teacher token, multiple next tokens may complete the same remaining bytes, but the teacher only specifies the required completion rather than how probability should be divided among these valid continuations. We introduce Event-Set Completion Distillation (ESCD), which complements cross-tokenizer probability alignment with completion-set supervision. ESCD aggregates prefix-related teacher events and supervises the total probability of byte-compatible one-step student completions, avoiding tokenizer-dependent probability splits among individual tokens. The method reuses student trajectories and predictions, requiring neither additional rollouts nor changes to the student vocabulary. Experiments demonstrate consistent gains in mathematics, code, and scientific reasoning across model families and tokenizers, extending to large-scale MoE distillation from a 1T teacher to a 35B student. Local analyses show that retaining completion sets better matches the reference supervision, while one-step completion covers over 99% of observed compatible teacher mass after partial event entry in the studied tokenizer pairs. These findings support event entry and event completion as complementary supervision targets for cross-tokenizer knowledge transfer. Code will be released on GitHub.