Continuous latent diffusion improves mathematical and code reasoning accuracy

Reasoning with Continuous Latent Diffusion

Artificial Intelligence

Summary

Solving complex problems like math questions or coding tasks often requires step-by-step thinking, which is hard for AI. This paper shows how to use a method called continuous latent diffusion to generate reasoning steps by gradually refining solutions inside a compressed space. The researchers found that just decoding answers well isn't enough; instead, using compact learned representations and special training tricks helps produce better reasoning and code generation. Their approach outperforms previous methods on challenging math problems and code writing tests.

What this means in practice

  • For software engineers: Improve automated code generation tools to write correct code snippets by applying more accurate reasoning through continuous latent diffusion models.
  • For data scientists: Enhance mathematical problem-solving AI systems used for data analysis by leveraging improved latent reasoning techniques.

Authors

Xiang Cheng

Abstract

Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/LFRM