A new method improves one-step AI image generation accuracy
Unifying Distributional Training for One-Step Visual Generation
Machine Learning
Summary
Generating realistic images in one step is hard because the system must match many complex details all at once. The authors propose a new approach called MGFlow that better captures and compares complex image features by grouping them into flexible clusters. This helps the model avoid common problems like focusing too narrowly on certain details and produces much better images on a standard big dataset. They also show it can improve text-to-image models to create good pictures faster.
What this means in practice
- •For image generation engineers: Improve one-step image generation models by using MGFlow’s Gaussian mixture feature modeling to reduce errors and mode collapse.
- •For text-to-image developers: Enhance text-to-image generators by post-training with MGFlow to deliver better quality images in fewer steps.
Authors
Chi Zhang, Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
Abstract
\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with \textbf{1.45} $\mathrm{FDr}^6$ on pMF-H and \textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/