Variational learning improves discrete diffusion for better sequence generation
Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion
Machine LearningArtificial Intelligence
Summary
Generating sequences like molecules, text, or playlists often needs making corrections as the sequence builds. The researchers developed a new approach that learns which mistakes to create during training to help the model get better at fixing them. This method treats the problem like a game between two parts: one decides how to corrupt data, and the other learns to fix it. Their approach improved the quality of molecules, text, and playlists generated, compared to older methods.
What this means in practice
- •For drug discovery teams: Improve molecular generation by increasing the validity of molecules produced through learned corruptions in discrete diffusion processes.$Commercial implications: Enables better molecular design tools that can generate valid drug-like molecules for pharmaceutical companies.
- •For natural language developers: Reduce perplexity in text generation by training denoisers with learned token substitutions rather than uniform or fixed masking.
Authors
Dmitrii Moor, Federico Tomasi, Paul N. Bennett, Alice Wang, Mounia Lalmas
Abstract
Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tokens once they are unmasked, while uniform diffusion permits revisions but relies on uniformly random token substitutions. We instead learn which substitutions are most useful for training the denoiser to re-draft. We introduce Variational Stackelberg Discrete Diffusion (VSDD), a framework for learning a semantically aware corruption process. VSDD formulates training as a leader-follower game: the leader defines a Markovian corruption process parameterized by the denoiser's token embeddings, while the follower optimizes a variational denoising objective with the corruption process held fixed. The leader rewards corruptions based on how much the denoiser improves after learning from them, rather than on how easily the current denoiser can reconstruct them. We measure this improvement under a fixed reference corruption process, approximate the follower's response with a one-step gradient update, and optimize the leader using a score-function estimator. We evaluate VSDD across molecular, text, and playlist generation. VSDD substantially improves molecular validity over uniform and masked diffusion, reduces text perplexity relative to uniform diffusion while remaining competitive with masked diffusion, and achieves sizable improvements in offline playlist recommendation metrics.