Proxy2World generates detailed scenes using simple scene proxies
Proxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing Them
Computer Vision and Pattern Recognition
Summary
Creating detailed videos of scenes usually needs complex guides showing exactly how the scene looks and moves, which is hard to gather automatically. This paper presents Proxy2World, a system that learns to generate realistic visuals from simple, rough versions of scenes and movements without needing paired data. It blends depth and color information to produce detailed images that follow the given scene structure. The authors also introduce ProxyBench, a way to test how well such systems work across different scenes and movements.
What this means in practice
- •For game developers: Generate immersive and controllable game environments from simple scene layouts without extensive manual design.
- •For film production teams: Create natural-looking visual effects and scene renderings guided by rough proxy data rather than detailed models.
Authors
Hongli Xu, Weilong Yan, Anbang Wang, Chunyu Zou, Siyu Hong, Jingwei Huang
Abstract
Lightweight scene proxies let creators control scene layout and motion while leaving room for imagination in appearance, lighting, and visual effects. However, a suitable proxy is not uniquely defined, making paired proxy-video data difficult to construct automatically at scale. We present Proxy2World, a controllable world model that learns these complementary capabilities from ordinary posed RGBD videos, without training on authored proxy-video pairs. The model jointly learns depth-conditioned RGB generation and joint RGBD generation through cross-modal flow matching. Learning both tasks enables proxy-camera hybrid denoising at inference to follow the proxy structure while producing natural, detailed visuals. We further introduce ProxyBench to evaluate this capability across a diverse set of scenes, camera trajectories, and subject motions. Experiments on ProxyBench show that Proxy2World achieves a better balance between structural adherence and visual quality than camera-controlled and geometry-conditioned methods, supported by quantitative metrics, VLM assessments, human evaluations and diverse qualitative results.