WorldWeave builds expandable 3D worlds for consistent video scenes
WorldWeave: Growing Persistent Geometric Worlds for Video Generation
Computer Vision and Pattern Recognition
Summary
It is hard for computer models to remember and keep track of 3D scenes when they get bigger or are seen from different angles. The authors present WorldWeave, a system that separates keeping track of the 3D world from making videos of it. WorldWeave creates maps of terrain, plans scenes carefully, and uses those maps to guide video rendering without changing the stored world data. This helps the system remember details over time and keep videos consistent, even when exploring new areas or revisiting old ones.
What this means in practice
- •For game developers: Create large, expandable game worlds that maintain consistent 3D structure as players explore or revisit areas.$Commercial implications: Enables selling games with seamlessly expanding 3D environments that remain consistent across player sessions.
- •For film and animation studios: Generate videos with coherent geometric scenes from varying camera paths while preserving structural consistency during editing.
Authors
Yifan Huang, Lifan Jiang, Qingyue Hao, Cheng Chen, Boxi Wu, Xiaoxue Ren, Xiaofei He, Dehai Zhao
Abstract
Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.