Video prediction and image generation improved by physics based model
Lagrangian--Hamiltonian Flows for Video Prediction and Image Generation: A Symplectic Perspective
Computer Vision and Pattern Recognition
Summary
Predicting future video frames or creating new images can be challenging because it requires understanding how images change over time. The authors use ideas from classical physics to represent images and their motion as mathematical flows, which helps model these dynamics more accurately and efficiently. Their method, called LHFM, uses these flows to predict videos and generate images with better accuracy and lower computational cost compared to some existing methods.
What this means in practice
- •For computer vision engineers: Generate accurate future video frames efficiently by integrating physics-inspired dynamic models.
- •For machine learning engineers: Improve image generation quality by utilizing physics-based flow matching techniques for better fidelity scores.
Authors
Jiawei Hu
Abstract
We introduce LHFM, a geometric framework for learning image dynamics. Drawing on structures central to classical mechanics, symplectic geometry, and geometric quantization, LHFM represents each image as an exact Lagrangian graph and models its evolution through image-dependent Hamiltonian flows, which yield a transport--source parameterization of image velocities. Our primary application is deterministic video prediction: LHFM-V is a recurrent model that advances frames by integrating predicted transport and source fields, and achieves the lowest reported FLOP count among the compared recurrent models with similar prediction accuracy. The image variant, LHFM-I, shows that the same construction is compatible with flow matching: in a matched experiment, it attains a lower FID than the flow-matching baseline.