Diffusion transformers learn faster by injecting encoder info

Scaffold Then Internalize: Representation Injection for Diffusion Transformers

Computer Vision and Pattern RecognitionMachine Learning

Summary

Training diffusion transformers to generate images or data can take a long time. The authors study a new way to speed this up by letting the model actively use information from a pretrained visual encoder during training, instead of just matching it afterward. Their approach, called REPI, first uses encoder data as a temporary guide and then helps the transformer learn to internalize it. Combining this with older methods leads to much faster training without losing quality.

What this means in practice

  • For machine learning engineers: Reduce training time for diffusion transformer models in image and data generation tasks by integrating encoder features more effectively.
  • For ai product developers: Enable faster iteration and deployment of diffusion-based generative models by cutting training costs and time significantly.$Commercial implications: Supports quicker production of generative AI products by reducing costly GPU training from millions to fewer steps.

Authors

Han Fu, Jiacheng Chen, Baoquan Zhao, Weidong Chen, Wei Liu, Qing Li, Xudong Mao

Abstract

Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce \textit{REPresentation Injection} (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over $43.5\times$. Code will be available at https://jeneveuxpas.github.io/REPI