ShotPlan: Cinematic Video Generation with Learnable Planning Token

2026-07-20Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors created ShotPlan, a new method to make longer movies from videos by planning individual scenes or shots more clearly. They use special tokens that help the model understand when one shot ends and another begins, working at the level of individual frames. This lets the model keep the story consistent across multiple shots better than before. Their tests show ShotPlan works better than older methods for making cinematic videos.

video diffusion modelscinematic video generationmulti-shot compositionshot planningplanning tokensFractional Temporal Rotary Position Embedding (FRoPE)frame-level modelinginter-shot consistency
Authors
Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang, Cong Liu, Junqi Liu, Haibin Huang, Hongxun Yao, Chi Zhang, Xuelong Li
Abstract
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.