Diffusion models first learn general features then memorize training data

First Learn, Then Memorize: The Spectral Bias of Diffusion Models

Machine Learning

Summary

Generating new images or data with diffusion models involves first learning general patterns and only later copying exact training examples. The authors found that this happens because of how the model’s training process breaks down into two parts, revealed by studying a matrix linked to the learning algorithm. The first part captures broad features and comes early in training, while the second part, linked to memorizing specific noisy versions of data, emerges later and takes much longer. They confirm this pattern both mathematically and through experiments with real image models.

What this means in practice

  • For machine learning engineers: Adjust training to delay memorization and improve generalization by modifying noise sampling or regularization in diffusion models.
  • For ai product teams: Control when diffusion-based image generators start copying training images too closely to maintain output novelty and quality.

Authors

Raphaël Urfin, Tony Bonnaire, Giulio Biroli, Marc Mézard

Abstract

Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample ($m$ noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for $m=1$. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size $n$. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear ($n \asymp d$) and polynomial ($n \asymp d^k$) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank $r$ tunes the generalization--memorization transition, and an $L_2$ penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.