4D Human-Scene Reconstruction from Low-Overlap Captures
2026-07-10 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors address the challenge of creating detailed 3D videos of people using only a few cameras that don't cover much area, which usually results in poor quality and missing parts. They introduce StudioRecon, a method that separates the background from people, improves the background view by generating many fake camera angles, and uses smart techniques to track and model the person’s movements accurately. Their approach also refines the final video to reduce errors. They show that their method works better than others on several real-world tests and can be used for things like making new camera paths or swapping people in videos.
volumetric capture4D reconstructionlow-overlap camerasvideo diffusion modelsdeformable Gaussian humansmulti-view keypoint fittingnovel view synthesistrajectory renderingbackground synthesis
Authors
Minhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, Jaesik Park
Abstract
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement.