Embedding prediction improves image generation quality in diffusion transformers
Embedding Prediction Helps Image Generation
Computer Vision and Pattern RecognitionMachine Learning
Summary
Generating images using artificial intelligence often relies on giving the AI a starting instruction like a word or category. Usually, this instruction stays the same as the AI progressively refines the image. The authors show that predicting these instructions anew at each step, based on how the image looks at that point, helps improve image quality. They build a model that predicts multiple instruction points simultaneously and then used these predictions to guide image creation dynamically. Their approach achieves better results on a standard image dataset while using less training effort compared to some previous methods.
What this means in practice
- •For computer vision engineers: Create higher-quality images efficiently by dynamically updating generation instructions during sampling in diffusion transformer models.
- •For content creation software developers: Integrate adaptive embedding conditioning into image generation tools for improved output quality with reduced training resources.$Commercial implications: This enables commercial imaging software to deliver better results faster, appealing to users needing efficient, high-quality generative image features.
Authors
Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu
Abstract
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.