Reverse diffusions contract divergences and guarantee local stationarity

First-Order Stationarity of Reverse Diffusions

Machine Learning

Summary

This paper studies how certain random processes called reverse diffusions behave when used to generate data. The authors show these reverse processes shrink a measure of difference between probability distributions at a steady exponential rate when the added noise has a particular mathematical property called strong convexity. They also analyze how discretizing these continuous processes still leads to stable results that locally match the correct data patterns. This helps understand how these methods produce realistic samples and relates closely to ideas in optimization.

What this means in practice

A theory result. No direct application yet.

Authors

Zhifeng Chen, Chenyang Jiang, Yazhen Wang

Abstract

Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse diffusion, absent in the reverse process based on ODEs. Second, we incorporate discretization and establish averaged first-order stationarity bounds---the sampling analog of averaged gradient-norm guarantees in nonconvex optimization---for samplers of both overdamped and underdamped diffusion models. As in nonconvex optimization, the convexity-free certificate is local: it guarantees score consistency, not global mode weights.