Papers for

film post-production teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

DynaTokens improve video models to show moving objects with camera motion

DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time

Abstract: Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Making videos that change when the camera moves is tricky because the movement comes from two places: the camera itself and things moving inside the scene. The authors found that existing video models do well when the scene is still but have trouble when things move. They created DynaTokens, small learnable pieces that help the model understand how objects move in a scene without changing the whole system. This approach lets the model show moving objects correctly even if the camera moves in new ways.
Open → 2609.35704v1

Ego forge creates first person videos from third person footage

Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation

Abstract: Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Seeing what a person sees just from videos taken by others is very hard because the viewpoint changes a lot and some things are never seen before. The authors developed Ego-Forge, a system that can guess what someone would see from their own eyes given third-person video, without needing extra instructions or special text captions. It learns from many videos and uses a clever new way to think about the scene without relying on complex geometry rules. This makes Ego-Forge faster, more general, and better at imagining new views than previous methods.
Open → 2609.35368v1