Inferring dynamic changes from distribution snapshots requires multiple time points
Beyond Gradient Flow: Identifiability and Recovery from Distribution Snapshots
Machine Learning
Summary
It is very hard to figure out how something changes over time if you only look at snapshots of where it is at different moments. The authors show that looking at just one snapshot doesn’t provide enough clues because some hidden details remain invisible. By comparing multiple snapshots over time, it becomes possible to uncover parts of the hidden dynamics. They also explain how to estimate these changes reliably using mathematical tools and provide limits on how well you can do this depending on the number of snapshots and samples. Their experiments confirm the theory and reveal trade-offs in the process.
What this means in practice
- •For statistical modelers: Infer underlying dynamic processes more accurately from multiple time-indexed distribution snapshots.
- •For data engineers: Design sampling schemes that collect snapshots at enough time points to reduce ambiguity in dynamic reconstruction.
Tested on simulated data.
Authors
Nam D. Nguyen, Valeriya Malysheva
Abstract
Inferring dynamics from snapshots of evolving distributions is fundamentally underdetermined: the Fokker-Planck equation constrains the drift $F$ only through its score-weighted divergence $\nabla\cdot F+F\cdot\nabla\logρ$, leaving a $ρ$-solenoidal gauge invisible to any single-time constraint. Time-indexed transport formulations cannot resolve this ambiguity: every admissible marginal path admits a curl-free explanation, minimum-action reconstruction selects it, and marginal fit alone cannot distinguish dynamically inequivalent explanations. Requiring one autonomous field to explain several marginals instead makes part of the hidden circulation visible as $\nabla\logρ$ changes across marginals. Separating instantaneous Fokker-Planck source constraints from the snapshot experiment, we show that the source constraints identify the field modulo the kernel of a stacked score-weighted divergence operator. For generic Gaussian shape variation, source constraints at $K\ge m$ time points in intrinsic dimension $m$ eliminate every polynomial gauge direction, whereas finitely many density snapshots alone admit aliasing; we give the obstruction explicitly. At a Gaussian anchor, for Sobolev smoothness $s$ and $n$ samples per time point, we derive a conditional lower rate $(nK)^{-2s/(2s+m+1)}$ for the tangent snapshot experiment, with a matching upper rate in a degreewise benchmark. Strong-form fitting is non-orthogonal to score error and cannot be repaired by spectral filtering. Instead, we estimate using smooth test functions while retaining the known diffusion term, and derive a finite-sample bound that separates sampling error from fixed-grid quadrature bias. Planted-circulation experiments confirm the predicted gauge contraction and expose a design tension between cross-slice information and covariance-aware whitening.