Video generation improved by layout control for large viewpoint changes
LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
Computer Vision and Pattern Recognition
Summary
Generating videos from a single image while controlling how the camera moves is hard, especially when the camera shows parts of the scene that weren’t visible before. The authors developed a method called LIFT that lets users specify not just the camera movement but also what objects should appear and where in a future view, particularly the last video frame. To train this, they created a special technique called on-policy self-distillation to teach the system to follow sparse layout instructions. They also built a dataset to test large viewpoint changes and found that LIFT produces better videos with more precise content control.
What this means in practice
- •For video game developers: Create realistic game cutscenes where the player controls both camera movement and scene content under large viewpoint shifts.$Commercial implications: Enhances interactive storytelling in games by allowing developers to generate controllable dynamic scenes, enabling richer user experiences.
- •For film and animation studios: Produce videos where directors specify final scene layouts along with camera paths, simplifying pre-visualization of complex shots.
Authors
Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu
Abstract
We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.