Multi-token prediction boosts language model GPU speed by nearly double
Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
Artificial IntelligencePerformance
Summary
Generating text one word at a time with language models on GPUs can be slow because the computer does many repeated steps. The authors studied predicting two words at once instead of one, then checked how this worked on a specific NVIDIA GPU. They found this method nearly doubled output speed and made the first words appear faster, even though the calculations got more complex. The speed gain came from doing fewer repeated runs of parts of the computation, balancing the extra work needed for guessing multiple words at once.
What this means in practice
- •For machine learning engineers: Speed up deployed language model inference systems on NVIDIA GPUs by using two-token prediction to reduce repeated GPU executions.
- •For gpu performance engineers: Optimize GPU workload management in language model inference by balancing complex execution paths against fewer kernel launches.
Authors
Suwesh Prasad Sah
Abstract
Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deployment on an NVIDIA A10G GPU. A 360-request benchmark covered plain-text, reasoning-intensive, and tool-calling workloads, while runtime telemetry, Nsight Systems, PyTorch Profiler, and selected Nsight Compute measurements were used to explain the observed performance. MTP increased output throughput by \(1.91\times\) to \(2.19\times\) across all prompts and reduced time to first output by 10.0--14.2\%. Median mean acceptance length ranged from 2.370 to 2.595 tokens per verification iteration. Profiling showed that MTP introduced a longer and more complex execution path, including proposal, sampling, attention, gathering, and reduction operations. However, it required 56.4--78.1\% fewer executions of the selected repeating CUDA Graph per generated token. The dominant MTP GEMM kernel was not faster than the dominant autoregressive GEMV kernel, and selected instances of both approached the A10G memory-bandwidth limit. These results show that MTP improved inference through amortization: greater token progress reduced repeated GPU execution sufficiently to outweigh the additional speculative-execution cost.