DominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decoding

2026-07-09Computation and Language

Computation and Language
AI summary

The authors introduce DominoTree, a new method to speed up large language model (LLM) text generation by drafting multiple tokens at once and checking them efficiently. Unlike earlier methods that treat token choices independently, DominoTree uses a correction that accounts for the order tokens appear in, improving accuracy without extra training. Their approach builds the token draft tree quickly on GPUs, maintaining the same output quality as slower methods. Tests on different models show DominoTree is faster and accepts longer token drafts compared to previous techniques.

large language modelsspeculative decodingtoken draftbest-first searchconditional correctionnon-factorized distributionGPU accelerationautoregressive decodingtop-M candidatestemperature (sampling)
Authors
Saw S. Lin, Jyh-Shing Roger Jang
Abstract
Speculative decoding accelerates LLM inference by drafting several tokens and verifying them in parallel. Block-diffusion drafters such as DFlash produce a draft block in one pass but model only per-position marginals; best-first tree methods such as DDTree expand candidate trees from those marginals. The released Domino drafter adds a GRU-based causal correction that makes each draft token's distribution path-dependent, a structure DDTree's factorized formulation cannot represent. We introduce DominoTree, a training-free best-first draft tree scored by Domino's conditional, non-factorized correction along each root-to-node path, made practical by restricting the per-node correction to a candidate top-M. On Qwen3-4B across eight benchmarks, DominoTree reaches up to 6.6x speedup over autoregressive decoding and the highest mean accept length of any evaluated method, up to 10.7 tokens per round, at every temperature we test. DominoTree constructs its tree with a GPU-native, CUDA-graph builder that is bit-identical to a reference Python implementation, so acceptance is unchanged, while keeping per-round tree construction cheap. With this builder as default, DominoTree wins throughput over the released Domino decoder at every temperature, 9-10% overall on Qwen3-4B and up to +22% on Alpaca, and over DDTree/CaDDTree at every temperature we test. On Qwen3- 8B, DominoTree keeps the highest accepted length at every temperature and adds a decisive throughput win at T=0, +24% over DDTree; at higher temperature that edge over DDTree/CaDDTree narrows to a tie and a small loss, while its Overall aggregate wins over DFlash and Domino persist.