GenNVS improves 3D view synthesis using geometry and diffusion models

GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Making new views of a scene from a single photo is hard because we don’t know the exact 3D shape of objects. The authors developed GenNVS, a system that uses a special 3D representation to separate foreground objects from the background and aligns them into one 3D scene. This helps a video diffusion model create new views that better keep the shapes and arrangement of things in the scene. Tests showed GenNVS produces clearer, more accurate images and allows easy changes to the scene afterward.

What this means in practice

  • For 3d artists and animators: Create consistent new views of scenes from single images while preserving object shapes for more realistic scene manipulation and editing.
  • For augmented reality developers: Generate accurate and spatially coherent 3D views from limited input to improve object integration and scene realism in AR applications.

Authors

Yajiao Xiong, Youyu Luan, Xiaoyu Zhou, Yongtao Wang

Abstract

Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Specifically, GenNVS models foreground objects and the background with 3D Gaussian Splatting and aligns them through a coarse-to-fine geometric optimization process to form a unified 3D scene. This scene conditions a video diffusion model through the proposed Dual-Stream Masking mechanism, which guides synthesis by jointly exploiting rendered validity masks and geometry-aware warping. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing.