Looped transformer improves text to image generation with fewer resources

Looped Diffusion Transformer

Computer Vision and Pattern RecognitionMachine Learning

Summary

Generating images from text usually requires bigger models or more processing steps. The authors show a way to reuse parts of a model multiple times during generation to improve image quality without increasing model size. They solve problems from naive reuse by adding extra learning signals and a special attention method to keep important details. Their method produces better images than bigger models while using less computing power and enables correcting earlier mistakes through repeated processing.

What this means in practice

  • For ai software engineers: Deploy efficient text-to-image models that produce higher quality images with fixed model size by using looped transformer blocks during inference.
  • For mobile app developers: Implement computationally cheaper text-to-image features on devices by replacing larger models with looped models that require less inference compute.

Authors

Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang

Abstract

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.