Summary
It is hard to control video scenes because moving a camera and objects in two dimensions can mean many different things in three dimensions. The authors created Generative Cinematographer (GenCine), a system that turns a single image into a 3D scene where artists can move the camera and parts of objects in 3D space. This system uses colored handles to let artists control how objects move piece by piece, even for complex shapes without needing special physics tools. GenCine then uses a special method to turn these 3D controls into instructions for a video generator, so the final videos follow the artist's motions realistically. It works on real and synthetic videos and keeps object shapes consistent when the camera changes viewpoint.
What this means in practice
- •For video game developers: Create animated 3D scenes where camera and object movements are jointly controlled with precision from a single image.
- •For film visual effects teams: Produce controlled special effects shots by editing camera and foreground motion in 3D while ensuring geometric consistency across viewpoints.
Authors
Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad
Abstract
Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.