Pixel space training improves few step text to image generation
DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses
Computer Vision and Pattern Recognition
Summary
Generating images from text usually takes many steps and complex systems working on compressed image data. The authors explore a way to speed this up by training models directly on normal RGB images (the pixels you see) instead of compressed forms. They find that matching textures at some noise level and using a new way to guide learning with existing visual representations helps create good images faster. Their new method, called DMA², produces high-quality images in fewer steps than previous models, making fast text-to-image creation more practical.
What this means in practice
- •For ai image generation developers: Create faster text-to-image generation models that produce high-quality RGB images in fewer steps, improving efficiency without sacrificing quality.
- •For computer graphics engineers: Integrate pixel-based training techniques into graphics pipelines to accelerate content generation from textual prompts while maintaining semantic fidelity.
Authors
Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Shaoteng Liu, Lehan Yang, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen
Abstract
Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA$^2$. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA$^2$ student performs better than the 25-step teacher and evaluated few-step distillers.