Unbalanced optimal transport improves one-step image generation quality
One-Step Generative Modeling via Unbalanced Optimal Transport
Information TheoryMachine Learning
Summary
Generating images in a single step is faster but can be tricky because the model tries to match two sets of images exactly, which can cause mistakes during training. The authors found that allowing the real images’ importance to change while keeping generated images fully accounted for leads to better picture creation. They designed a method called Unbalanced Optimal Transport Gradient Flow (UOT-GF) that uses this idea and showed it improves image quality on ImageNet. They also studied the math behind this method to understand why it works and when it succeeds.
What this means in practice
- •For image generation developers: Build faster one-step image generators with improved quality by using unbalanced transport to better match training data distributions.
- •For machine learning engineers: Improve model robustness and convergence in large-scale training of generative models by applying asymmetric mass matching in optimal transport.
Authors
Yirong Shen, Mengfei Xia, Junpeng Jing, Lu Gan, Cong Ling
Abstract
Drifting models enable one-step generation by amortizing distribution transport into training, but this efficiency places greater demands on the transport field estimated at each update. In large-scale training, the field is computed from finite mini-batches of generated and real samples, which provide only imperfect approximations to the underlying distributions. Balanced optimal transport enforces exact mass matching within every mini-batch, making the estimated field sensitive to the particular composition of the real-data batch. We find that generated and real samples should be treated asymmetrically: letting the mass assigned to real samples adapt while keeping every generated sample fully transported improves generation across six feature-space metrics in controlled ablations, and is more robust to the relaxation strength than relaxing both marginals simultaneously, which falls below balanced transport under stronger relaxation. Motivated by this observation, we propose Unbalanced Optimal Transport Gradient Flow (UOT-GF), which keeps the generated-sample marginal fixed and relaxes only the real-data marginal. Under identical settings at DiT-B/2 on ImageNet-256, UOT-GF improves Fréchet Inception Distance (FID) from 1.53 to 1.46 over the balanced W-Flow baseline; scaling the same recipe yields 1.34 and 1.22 FID at L/2 and XL/2, the best FID among the one-step models we compare. We further derive the induced UOT transport force, establish a kinetic Vlasov--Fokker--Planck formulation whose overdamped zero-temperature limit recovers the drifting dynamics, and characterize non-target stationary states together with sufficient conditions for convergence.