Looped transformers improve accuracy scaling with adaptive iterations
Improving Test-Time Scaling with Adaptive Looped Transformers
Computation and LanguageMachine Learning
Summary
Longer outputs in transformer models require more computing to stay accurate. The authors studied looped transformers that reuse the same layers multiple times to save parameters but noticed many tokens don’t benefit equally from repeated processing. They created TaH2, which adaptively decides which tokens need extra attention, improving accuracy without extra costs. This method made models more efficient and precise especially as output length and iteration depth increase.
What this means in practice
- •For natural language processing engineers: Increase efficiency and accuracy of transformer-based models during text generation by focusing computational effort on important tokens at test time.
- •For speech recognition developers: Improve transcription accuracy in speech-to-text systems by adaptively allocating processing iterations during decoding based on token difficulty.
Authors
Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang
Abstract
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.