Block verification aware loss improves token acceptance in speculative decoding
BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
Artificial IntelligenceComputation and Language
Summary
When computers try to predict sequences of words or code, they often check one word at a time to make sure predictions are correct. This paper introduces a new way to train these systems that focuses on checking whole blocks of words together instead of one by one. The new method helps the computer accept longer chunks of predictions at once, making the process faster and more efficient. The authors show improvements across tasks involving math, coding, and chatting without changing how the computer checks its predictions during use.
What this means in practice
- •For natural language processing teams: Train language models to generate sequences more efficiently by increasing accepted token blocks during decoding.
- •For software engineers building coding assistants: Improve code generation speed and accuracy by using loss functions aligned with block-level verification in diffusion-based models.
Authors
Suyoung Kim, Jahyun Koo, Hyeonjin Kim, Inhyeok Bang, Seunghyun Lee, Hyunjae Oh, Baeseong Park, Dongsoo Lee
Abstract
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0--21.0\% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.