Compute allocation improves training quality in drifting image models

Refresh or Realize? Compute Allocation in Drifting Models

Machine Learning

Summary

This paper looks at how to best use computing power when training drifting models, a type of image generator. The models learn by moving images step-by-step based on calculated directions, but the directions change over time, so the model must decide whether to spend more time refining current directions or recalculating new ones. The authors found that early on, more focus on refining helps, but later, recalculating directions more often leads to better image quality. This insight helps balance computation during training to get better final images.

What this means in practice

  • For image generation developers: Optimize training schedules of image generators by balancing parameter updates and frequent recalculation of motion targets for better quality outputs.
  • For machine learning engineers: Design compute budgets for training models that drift over time by prioritizing fresh target computations to improve efficiency and output quality.

Tested on one dataset.

Authors

Sipeng Chen, Xu Zheng, Shibo Li

Abstract

Drifting Models train a one-step generator by recomputing a finite-sample drift field at every iteration and taking an optimizer step toward the drifted target. The field says how generated samples should move, but the step is taken in parameters shared by all samples, so the motion the network actually makes need not match the motion it was given. This leaves a basic training question open: should extra compute go into fitting the current target more closely, or into recomputing the field? We study it on ImageNet 256x256. Holding the target fixed for k optimizer steps and measuring the realized displacement, we find that deeper fitting does bring the network closer to the frozen target, and that the number of steps needed before it makes any net progress drops from about sixteen early in training to one later on. When the extra steps come for free, k=2 also lowers FID. Once they are paid for, the result flips: at approximately matched measured wall-clock, spending the budget on fresh fields gives lower FID than deeper fitting, on both training seeds. The target itself shows why a fresh field is worth so much. Redrawing the finite support rotates its direction far more than a parameter update does (cosine ~0.3-0.6 against ~0.95), and a correction that is optimal in field space is not reliably better in FID than a parameter-free one. For Drifting, fitting each target well and spending compute well are different goals.