UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

2026-08-10Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors address challenges in editing 3D human motions using text instructions by creating a large, diverse dataset called Omni-MoEdit and designing a new model named UniMoFlow that improves how motions are generated and edited together. They also introduce a method called SAFE to refine edits without losing important original motion details. Their approach better matches the intended text description while keeping changes meaningful and consistent. They evaluate their method with new metrics that recognize acceptable variations beyond exact matches.

3D human motiontext-to-motion generationlatent flow matchingspatiotemporal localizationsemantic groundingmotion editingcycle consistencysource fidelitydataset synthesisgenerative models
Authors
Yilei Hua, Beibei Jing, Ce Zheng, Hanyu Zhou, Yawei Luo, Wei Yang
Abstract
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.