HiPhy improves video generation with realistic physical interactions

HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

Computer Vision and Pattern Recognition

Summary

Generating videos that follow real-world physics is hard, especially when multiple physical effects happen together. The authors propose HiPhy, a new method that makes videos where different physical rules work correctly at the same time, like a balloon floating while steam rises. They created a large dataset and benchmark to test this idea. Their method does better than previous ones at making videos that look physically and semantically believable.

What this means in practice

  • For visual effects studios: Generate video scenes where multiple realistic physical phenomena interact coherently, improving special effects authenticity in film and animation.$Commercial implications: Enables production of physically accurate video effects that can be commercialized for movies and games with complex physical interactions.
  • For robotics simulation teams: Create video-based simulations that better model multiple physical principles simultaneously for testing and training robotic systems.

Authors

Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag

Abstract

Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.