Triangular resampling improves long-term motion generation accuracy
Triangular Resampling for Long-Horizon Motion Generation
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Generating long sequences of human motion is hard because small mistakes add up over time. The authors propose a method called Triangular Resampling that helps a type of AI model stay accurate when creating longer motions. Their method works by mixing correct steps with the model’s own predictions during training to better handle long sequences. This approach was tested on motions lasting two minutes and showed significant improvements compared to previous methods.
What this means in practice
- •For animation studios: Generate longer and more accurate human motion sequences for digital characters using improved diffusion models.$Commercial implications: Enables production-quality animation tools to create complex motion over extended time frames with less error accumulation.
- •For virtual reality developers: Create more realistic and stable long-duration avatar movements by reducing error drift in motion generation models.
Authors
Kunhang Li, Yiyi Cai, Xiangyue Zhang, Fangyuan Tu, Yuhan Wu, Zhixiang Wang, Kaipeng Zhang, Haiyang Liu
Abstract
We introduce Triangular Resampling (TR), a post-training method for mitigating long-horizon error accumulation in motion diffusion models. Built on FloodDiffusion's triangular denoising schedule, TR addresses the mismatch between ground-truth-derived training windows and model-generated inference states. Replacing only completed motion history leaves this mismatch unresolved in partially denoised states within the active window. TR therefore extends rollout-based training to these states, using ground-truth clamping to limit excessive drift. For each replayed sample, TR draws one denoising threshold, shared across latent positions and replay updates, and replays multi-step triangular denoising without gradient tracking. After each update, states below the threshold are replaced with noise-matched ground truth, while those at or above it retain model predictions. The resulting latent window enters the standard training update. This rollout construction supports both supervised training (TR) and distribution matching (TR-DMD). On 120-second motion generation from HumanML3D test prompts, TR and TR-DMD achieve state-of-the-art FID AUC within their respective non-DMD and DMD comparison groups. Supervised TR reduces FID AUC by 40.9% and FID degradation slope by 55.3% relative to matched post-training without replay.