Diffusion previews speed up image selection with fewer steps

Unlocking Few-Step Diffusion for Faithful Previews

Machine LearningComputer Vision and Pattern Recognition

Summary

When people make images using AI that draws step-by-step, it usually takes a lot of steps and time. The authors found a way to make quick preview images in just a few steps that look much like the full images, by fixing the starting noise and how the AI cleans up the image. This helps users quickly see many options without waiting long, then only make the final images for the best ones. Their method works with different image generators and doesn’t need to be retrained for different speeds.

What this means in practice

  • For ai product developers: Create faster user previews in AI image generation apps by generating high-quality images using fewer steps before full rendering.$Commercial implications: Enables faster preview features in commercial AI art tools, improving user experience and reducing compute costs.
  • For cloud service engineers: Reduce resource usage by generating quick, faithful preview images on cloud platforms before committing to full-resolution generation.

Authors

Jing Jia, Sifan Liu, Guanyang Wang

Abstract

Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.