Representation alignment improves training speed of visual AI models

What Visual Generators Need from Teachers: Rethinking Representation Alignment

Computer Vision and Pattern Recognition

Summary

Training AI models that generate images is faster when parts of the model learn to copy specific layers of a well-trained reference model, called a teacher. The authors found that it only helps to copy teacher layers that the new model struggles to mimic on its own. They created a method called RARE that intelligently picks which teacher layer to copy and stops copying when improvements slow, saving time and producing better image quality. Their approach beats several older methods on standard benchmarks with less computing effort.

What this means in practice

  • For ai model training teams: Speed up training of image generation models by choosing optimal representation alignment strategies to improve quality with fewer resources.
  • For machine learning engineers: Reduce compute costs during vision transformer training by applying the RARE method to align teacher features selectively and phase out losses dynamically.

Authors

Yongcong Wang, Hingchin Chen, Mingyu Fan, Shuo Jiang, Teer Zhang, Yucong Sun, Zijia Wang, Yiming Lu, Chengchao Shen

Abstract

Representation alignment speeds up diffusion transformer training by pulling an intermediate block of the model (student) toward features of a frozen pretrained encoder (teacher). Which teacher layer to align, and for how long, is still set by convention, and each alternative costs a training run. We find that alignment helps where the student cannot linearly recover the teacher's features, not where it already resembles them. Since a deep teacher layer is largely predictable from the one below, we isolate what each layer adds, its increment, and measure how much of it an unaligned student recovers. The student fills the teacher's hierarchy from the bottom up and stalls near the top, which we call hierarchy filling: even after 400K steps it recovers almost none of the deepest. The recoverability gap is the unrecovered share of an increment, read from one unaligned checkpoint. In short runs that each align one teacher layer at one block, the gap nearly reproduces their ranking by FID improvement, and CKA, a measure of feature similarity, largely reverses it. Representation Alignment and Recoverability Estimation (RARE) picks the teacher layer with the largest gap before training. During training, it tracks each token's remaining distance to that layer, the online counterpart of the gap, weights tokens by it, and phases out the loss once the average distance stops falling. With SiT-B/2 on ImageNet $256\times256$, RARE reaches an FID of 18.02 without guidance and 4.46 with it, ahead of seven alignment baselines including REPA, iREPA and HASTE. It also trains in 14% fewer GPU-hours than iREPA. Its FID stays below iREPA's across model scales, teachers, datasets and backbones.