Video prediction and image generation improved by physics based model

Lagrangian--Hamiltonian Flows for Video Prediction and Image Generation: A Symplectic Perspective

Computer Vision and Pattern Recognition

Summary

Predicting future video frames or creating new images can be challenging because it requires understanding how images change over time. The authors use ideas from classical physics to represent images and their motion as mathematical flows, which helps model these dynamics more accurately and efficiently. Their method, called LHFM, uses these flows to predict videos and generate images with better accuracy and lower computational cost compared to some existing methods.

What this means in practice

Authors

Jiawei Hu

Abstract

We introduce LHFM, a geometric framework for learning image dynamics. Drawing on structures central to classical mechanics, symplectic geometry, and geometric quantization, LHFM represents each image as an exact Lagrangian graph and models its evolution through image-dependent Hamiltonian flows, which yield a transport--source parameterization of image velocities. Our primary application is deterministic video prediction: LHFM-V is a recurrent model that advances frames by integrating predicted transport and source fields, and achieves the lowest reported FLOP count among the compared recurrent models with similar prediction accuracy. The image variant, LHFM-I, shows that the same construction is compatible with flow matching: in a matched experiment, it attains a lower FID than the flow-matching baseline.