Weighting schedules control learning phases in generative models for mixed data
Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data
Machine Learning
Summary
Generative models create new data by learning patterns over time, but when learning data with multiple types or groups, they focus on each group at different times. The authors found that the way the model’s training is weighted over time determines when it learns the shapes of the groups and when it learns how frequently each group appears. They showed this by studying simple but high-dimensional examples and confirmed it on images and genetic data. This understanding can help design better training methods for complex data.
What this means in practice
- •For machine learning engineers: Tune the weighting schedules during training of generative models to improve learning of complex data features and mode frequencies.
- •For bioinformatics teams: Enhance generative models for genetic data synthesis by adjusting training to capture population diversity and genotype frequencies more accurately.
Authors
Jérémie Klinger, Raphaël Urfin, Giulio Biroli, Marylou Gabrié
Abstract
Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time $t$ with a weighting schedule $w(t)$. Along the backward dynamics, and for multi-modal distributions, trajectories commit to modes of the target within a narrow time window, the \textit{speciation time}. In this work, focusing on high-dimensional data, we decompose the integrated loss into its single-time contributions and analyze each at fixed signal-to-noise ratio $Λ(t)$: we show that $Λ(t)$ sets the rate at which each feature of a multimodal target - the mode directions and their relative weights - is acquired during training. Crucially, at high $Λ(t)$ all mode directions are acquired together, on a single timescale insensitive to their amplitudes, while the relative weights are not learned at all. Only near the speciation time, where $Λ(t)$ becomes of order one, do all features become learnable, each on its own timescale: the weights are acquired jointly with the directions, and the directions at rates set by their relative amplitudes. For models trained on time-integrated objectives, the learning dynamics is then governed by how much of the weighting effectively sits near the speciation time, which provides insights on $w(t)$ design choices. These results follow from an exact high-dimensional analysis of the training dynamics of unbalanced and hierarchical Gaussian mixtures. Numerical experiments on image and human genome haplotype generation recover the predicted hierarchy of learning timescales in more complex settings.