Papers for

content creation software developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Embedding prediction improves image generation quality in diffusion transformers

Embedding Prediction Helps Image Generation

Abstract: In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

Thu 1 OctComputer Vision and Pattern RecognitionMachine Learning
The gist
Generating images using artificial intelligence often relies on giving the AI a starting instruction like a word or category. Usually, this instruction stays the same as the AI progressively refines the image. The authors show that predicting these instructions anew at each step, based on how the image looks at that point, helps improve image quality. They build a model that predicts multiple instruction points simultaneously and then used these predictions to guide image creation dynamically. Their approach achieves better results on a standard image dataset while using less training effort compared to some previous methods.
Open → 2610.02203v1