Physis-lang improves physical accuracy in video prediction models
Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model
Computer Vision and Pattern Recognition
Summary
Video models often create videos that look real but don't follow the rules of physics. The authors show that using detailed physical descriptions in language can help these models understand how things really behave. They built a system called Physis-Lang that creates and improves these descriptions, helping models make more physically correct videos. Tests show these improvements work across different video datasets and models.
What this means in practice
- •For video game developers: Build game simulations with more physically accurate video scenes using language-enhanced model training.
- •For robotics engineers: Improve robot vision systems by training models that better predict realistic physical changes in their environment.
Authors
Liming Lu, Xianzheng Ma, Wenkun He, Guanqi Zhan, Yilin Zhao, Junyu Chen, Mengyao Xu, Jiaojiao Fan, Wenhang Ge, Yuchao Gu, Yunze Liu, Boyi Li, Zhen Dong, Victor Prisacariu, Ming-Yu Liu, Song Han, Han Cai
Abstract
Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.