Text to video system generates physically realistic complex motions
CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation
Computer Vision and Pattern Recognition
Summary
Videos created from text prompts often look smooth but don’t follow real-world physics. The authors designed a system called CompAdapt that can create videos respecting complex physical motions like collisions and multi-stage actions. It turns language instructions into detailed motion rules and adapts quickly to new kinds of physical behavior without re-training. Their experiments show this method makes videos that look better and act more physically correct than earlier systems.
What this means in practice
- •For visual effects teams: Generate realistic video scenes with complex physical interactions from text descriptions without manual physics tuning.
- •For game developers: Create adaptive video sequences with physically plausible motions for in-game cutscenes directed by natural language prompts.
Authors
Haoran Qin, Renlong Wu, Tianyu Huang, Yukang Ding, Hui Li, Wangmeng Zuo
Abstract
While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a physics-consistent T2V framework for adaptable generation across complex real-world scenarios. It extends neural dynamics modeling beyond single-type motions to encompass composite physical behaviors, including coupled motions, multi-stage transitions, and multi-object collisions. Furthermore, CompAdapt translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial physical parameters. To generalize to novel physical environments, CompAdapt introduces dynamics-aware prior matching, achieving one-shot adaptation without retraining the core dynamics module. In addition, a physics-aware latent feature fusion module improves visual fidelity under fast and complex motion. Experiments on physics-focused T2V benchmarks demonstrate that CompAdapt improves physical consistency over both general T2V models and physics-constrained baselines, while preserving high visual quality and adaptability to unseen dynamics. The project page is available at https://makapic.github.io/CompAdapt/ .