Bidirectional motion and text generation improves language and movement modeling
BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion
Computer Vision and Pattern Recognition
Summary
Generating text descriptions from human motion and creating motion from text are both important but challenging tasks, especially when the two need to work together. The authors propose a new method called BiMoGen that models these tasks in both directions simultaneously, using a way to predict parts of a sequence in any order to reduce errors that accumulate with one-direction models. They also add training tricks that help the model learn better and correct its own mistakes during generation. Testing showed that BiMoGen performs well on standard motion-text datasets, improving how well motion and text correspond to each other.
What this means in practice
- •For animation developers: Create more accurate motion sequences from textual descriptions to enhance character animation workflows.
- •For game designers: Generate textual summaries of game character motions to improve game content indexing and accessibility.
Authors
Wanjiang Weng, Yongliang Wu, Xiaofeng Tan, Xingyu Zhu, Wenbo Zhu, Hongsong Wang
Abstract
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs. The project page is available at https://wengwanjiang.github.io/BiMoGen-Page.