Looped transformers speed up with new depth asynchronous self speculation
Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers
Machine LearningArtificial Intelligence
Summary
Generating text with some repeating neural networks can be slow because each new word needs many step-by-step calculations. The authors found that certain parts of these networks finish learning earlier than others, which can be used to guess words quickly. They created a new method that lets early guesses use fully finished info without repeating all steps, making text generation 4 to 7 times faster. This could help make sophisticated language models respond faster without extra computing effort.
What this means in practice
- •For natural language processing engineers: Speed up generation of text from large transformer models by using asynchronous self-speculation on intermediate layers.
- •For software development teams: Enable faster code completion tools by reducing the latency of deep recurrent transformer decoders.
Authors
Guanghao Li, Zihan Su, Hao Yu, Jinyang Jiang, Tao Ren, Zehao Li, Feng Lu, Ming Tang, Chun Yuan
Abstract
Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce Depth-Asynchronous Self-Speculation (DAS), which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its Mature-V primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop DAS-Wave, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves 4.00--6.96$\times$ mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.