DynaTokens improve video models to show moving objects with camera motion
DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time
Computer Vision and Pattern Recognition
Summary
Making videos that change when the camera moves is tricky because the movement comes from two places: the camera itself and things moving inside the scene. The authors found that existing video models do well when the scene is still but have trouble when things move. They created DynaTokens, small learnable pieces that help the model understand how objects move in a scene without changing the whole system. This approach lets the model show moving objects correctly even if the camera moves in new ways.
What this means in practice
- •For video game developers: Generate realistic moving scenes with user-controlled camera paths and dynamic objects for immersive gameplay visuals.$Commercial implications: Improves dynamic scene rendering in games by enabling better integration of camera and object motion, enhancing player experience.
- •For film post-production teams: Create or edit video scenes where camera moves and object motions combine realistically without full model retraining.
Authors
Ma Ziqi, Chen Hongqiao, Gkioxari Georgia
Abstract
Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/