FloodDiffusion 2 speeds up and controls streaming motion generation
FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation
Computer Vision and Pattern Recognition
Summary
Generating smooth and natural human-like motion on computers can take a long time and be hard to control precisely. The authors improved an existing system, FloodDiffusion, by redesigning how it remembers past movements and adding a way to guide the character’s path exactly. Their method makes generating motion much faster—over ten times quicker during key steps—and produces better, more realistic movements. This helps create smoother animations for long sequences, useful in games or virtual environments.
What this means in practice
- •For game developers: Generate real-time or long-duration character animations with precise path control and efficient computation for interactive experiences.
- •For virtual reality designers: Create natural human motion in VR avatars that follow exact movement paths without costly computation delays.
Authors
Yiyi Cai, Yuhan Wu, Kunhang Li, Tu Fangyuan, Xiangyue Zhang, Qiaoge Li, Zhixiang Wang, Kaipeng Zhang, Haiyang Liu
Abstract
We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming motion generation model. While FD1 produces plausible motion, it suffers from low efficiency and limited controllability, as its attention design requires repeated computation over the entire history, and it lacks precise trajectory control for real-world applications. To address these limitations and improve generation quality, FD2 introduces three advances. First, Partial Attention makes finalized history representations independent of the active window, enabling KV-cached inference and shared-history packing for efficient training. Second, we establish a necessary-and-sufficient Bregman criterion for regression losses to preserve diffusion's conditional-mean velocity field. This criterion guides an FK-induced quadratic loss that incorporates motion geometry without online FK evaluation. Third, FD2 introduces precise path conditioning to control the character's root trajectory while preserving natural body motion. Experiments show that FD2 reduces training computation by 4.6$\times$ and accelerates denoising by 11.29$\times$, reaching 2.303 ms per update on long sequences. Alongside these efficiency gains, FD2 improves motion quality over FD1 and achieves state-of-the-art FID scores among streaming methods, with 0.048 on SEED and 0.053 on HumanML3D.