Hierarchical continuous diffusion models improve puzzle and language tasks

Hierarchical Continuous Diffusion Language Models

Computation and LanguageArtificial IntelligenceMachine Learning

Summary

Generating sentences or solving puzzles with computers is tricky because words depend on each other. The authors found a way to combine two methods: discrete words and flowing hidden signals, so the computer can consider the whole sentence at once while creating it step by step. This method improves accuracy in solving puzzles like Sudoku and planning problems, as well as making better guesses for the next words in text. Their approach keeps the hidden signals as the main state, reading and updating words from it continuously. This helps the computer keep words related and consistent while generating text or solving problems.

What this means in practice

  • For software developers: Improve AI systems that solve puzzles and complex planning tasks by using a model that better maintains token dependencies during generation.
  • For natural language processing teams: Enhance language generation systems with a model that reduces errors in predicting text by integrating hierarchical continuous diffusion.

Authors

Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing

Abstract

Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.