Looped transformers speed up with new depth asynchronous self speculation

Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers

Machine LearningArtificial Intelligence

Summary

Generating text with some repeating neural networks can be slow because each new word needs many step-by-step calculations. The authors found that certain parts of these networks finish learning earlier than others, which can be used to guess words quickly. They created a new method that lets early guesses use fully finished info without repeating all steps, making text generation 4 to 7 times faster. This could help make sophisticated language models respond faster without extra computing effort.

What this means in practice

Authors

Guanghao Li, Zihan Su, Hao Yu, Jinyang Jiang, Tao Ren, Zehao Li, Feng Lu, Ming Tang, Chun Yuan

Abstract

Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce Depth-Asynchronous Self-Speculation (DAS), which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its Mature-V primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop DAS-Wave, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves 4.00--6.96$\times$ mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.