Loop dropout improves adaptation of repeated transformer blocks in language models
Loop Dropout: Regularizing Shared Updates in Looped Language Models
Machine LearningComputation and Language
Summary
Looped language models use the same transformer block multiple times to process information, but adapting these repeated blocks evenly is hard. The authors found that a common technique, LoRA, works better near the end of these loops than at the start. They created Loop Dropout, a method that randomly skips some updates during training but keeps their overall strength balanced. This helps the model adapt more evenly across all loop positions and improves tasks like mathematical reasoning and code generation. The method does not add extra complexity during actual use and works better than previous approaches.
What this means in practice
- •For language model engineers: Enhance training of looped language models by improving adapter updates to boost reasoning performance without increasing inference cost.
- •For software developers: Improve instruction-following and code generation capabilities in AI-based coding assistants by using Loop Dropout during model fine-tuning.
Authors
Zirui Zhu, Hailun Xu, Xuanlei Zhao, Yong Liu, Yingxuan Ren, Kanchan Sarkar, Kun Xu, Yang You
Abstract
Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update is more effective at later loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation.