Iterative refinement improves dynamic 4D scene synthesis from sparse videos

4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors

Computer Vision and Pattern Recognition

Summary

Creating detailed 3D videos from just a few camera views is difficult because it’s hard to guess parts of the scene that aren’t seen. The authors developed a method that first builds a fuller 3D shape from sparse views, then uses AI to improve the video frames and fix mistakes step by step. This approach helps fill in missing details and makes the dynamic 3D scenes much clearer and more complete. Their method showed significantly better results than earlier techniques on a standard test.

What this means in practice

  • For virtual reality developers: Produce higher quality dynamic 3D environments from limited video inputs for immersive VR experiences.
  • For visual effects teams: Generate improved 4D dynamic scene reconstructions from sparse-view footage to enhance movie and game content creation.

Authors

Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang, Sheila Lim, Huangying Zhan, Yi Xu, Anyi Rao, Frank Guan

Abstract

This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.