New loss function improves mode coverage in flow-based samplers
Mode Coverage in Normalizing Flow Boltzmann Generators via Log-Ratio Variation
Machine Learning
Summary
Sampling complex systems often misses important parts called modes, leading to incomplete results. The authors introduce a new mathematical tool called the log-ratio variation to better capture these modes during training. This tool enhances the way models learn to represent all important parts of the target system. Their experiments show improved coverage and more reliable predictions compared to earlier methods, meaning the new approach detects and includes more important scenarios.
What this means in practice
- •For machine learning engineers: Train flow-based generative models to better capture diverse data modes by incorporating log-ratio variations for improved sampling reliability.
- •For computational chemists: Use improved Boltzmann generators to generate more comprehensive molecular configurations for simulations, reducing missing important states.
Authors
Qi Feng, Rongjie Lai, Di Qi, Xuda Ye
Abstract
Normalizing flow Boltzmann generators retain a tractable pushforward density, but training with forward KL depends on target samples that may be biased or omit modes. As a result, a flow can miss target mass while its observed importance weights give a high effective sample size. We introduce the log-ratio variation $\X_ω$, the mean absolute pairwise difference of the target-to-pushforward log-density ratio under a weighting measure $ω$, and use it to define KLXX, a new loss function. Two log-ratio variations are added to the forward KL (denoted by the two X's): one weighted by the target to improve accuracy, the other by a mixture of quench and temper samples with pushforward samples to search candidate modes. We derive the Fisher--Rao gradient flow of KLXX, where both variations contribute nonpositive dissipation, and a fixed-surrogate error bound for KLXX. We use KLXX in an adaptive-staging Boltzmann generator, with importance reweighting at every stage. We bound the sampling error of its inference scheme when the stage weights are essentially bounded, and prove it asymptotically unbiased in the sample size. In the numerical tests, KLXX improves mode coverage over forward KL. It also improves the generator's per-stage diagnostics against the loss that built the schedule. The observables the generator recovers are close to independent references. The log-ratio variations thus supply information that the forward KL loss usually omits.