Geometry directs attention for long video camera control

Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

Computer Vision and Pattern Recognition

Summary

Creating long videos controlled by a moving camera is hard because the system needs to remember and find old parts of the scene accurately. The authors found a new way to use geometric information not to rebuild the whole scene but to point attention to the right bits of visual memory. Their method, called GEAR, links parts of past frames to current views and filters out parts hidden by objects, improving video quality and camera control. This approach helps generate longer videos that keep consistent views even along tricky camera paths.

What this means in practice

  • For film production teams: Generate complex, consistent long camera shots in virtual scenes that recall previous visual information precisely for better visuals.$Commercial implications: Enables advanced camera-controlled video generation for virtual production systems that improve scene continuity and viewer immersion.
  • For game developers: Improve in-game cinematic sequences with camera movements that maintain clear and consistent scene details based on past views.

Authors

Zesong Yang, Weikai Chen, Liyuan Cui, Lutao Jiang, Runze Zhang, Yingda Yin, Xiaoyang Huang, Kai Yan, Keyang Luo, Wangguandong Zheng, Xin Wang, Hujun Bao, Zhaopeng Cui

Abstract

Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric errors. Our key insight is that geometry need not explain the scene--it only needs to determine where visual memory should be read from, while attention decides what should be recovered. Based on this insight, we introduce GEAR, a Geometry-Enabled Attention Routing framework that uses geometry as an explicit token-level address for visual memory. Rather than fusing historical observations into a persistent global 3D representation, GEAR retains them as frame latents and uses per-frame geometry only to establish token-level correspondences with target views, thereby avoiding persistent error accumulation from global fusion. Guided by these correspondences, Geometric Correspondence Attention (GCA) selectively injects geometrically matched historical features into noisy target patches during denoising. We further introduce an Invisible Octree to accumulate visibility evidence and reject geometrically plausible but occluded correspondences. Extensive experiments demonstrate that GEAR achieves state-of-the-art visual quality, precise camera control, and revisit consistency, enabling minute-long video generation along challenging trajectories.