Block diffusion improves language model training and speeds generation
d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
Machine LearningArtificial Intelligence
Summary
Large language models usually write text one word at a time, which can be slow. Some newer models write several words at once in blocks, making the process faster. The authors found that teaching these block models using a method aware of the future words they can see (called d-OPD) gives better results and speeds up training. This method aligns the teaching signals with what the block model actually knows while generating text. They tested this on models of various sizes and saw clear improvements.
What this means in practice
- •For natural language processing engineers: Train block diffusion language models faster with more accurate supervision matching the model’s context window.
- •For software developers making chatbots: Deploy language models that generate multiple words in parallel to reduce user wait time during conversations.
Authors
Ruitao Liu, Qinghao Hu, Song Han
Abstract
Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy distillation (OPD) has been widely used for LLM training because it supervises the student on states generated by its current policy, rather than only on fixed offline trajectories. By training on the states the student actually visits, it reduces the mismatch between training and generation and can provide more relevant supervision as the student evolves. Recent work has extended this idea to AR-to-block-diffusion conversion. However, this setting introduces a fundamental mismatch in supervision: the block-diffusion student and the causal AR teacher condition on different information at the same training state. The student predicts from the entire partially denoised block, including visible future context, whereas the standard AR teacher target is defined only from the causal prefix. As a result, the teacher distribution used for distillation is not fully aligned with the information available to the student. We therefore introduce d-OPD, a future-aware on-policy distillation method that corrects the AR teacher distribution to better align with the student-visible state by incorporating visible future information within each block, providing supervision that better matches the information used by the student. Across Qwen3 models from 0.6B to 8B, d-OPD improves the six-benchmark average by up to $4.0$ points over OPDLM and reduces training time by $1.35$-$1.58\times$. The code is available at https://github.com/mit-han-lab/d-OPD.