NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction

2026-08-24Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors developed NemoSplat, a new method to create detailed 3D scenes from underwater videos, which are usually hard to process because water distorts light and moving objects cause problems. Their system improves camera tracking and scene depth understanding by separating moving things from the background, using smart guesses and even optional text hints to know what is moving. They also made a way to fix the visual problems caused by water, so the final images look clearer. To test their method, they created a big underwater video dataset with lots of movement and showed that their approach works better than existing ones.

underwater imaginglight scatteringnovel view synthesis4D Gaussian Splattingdynamic scene reconstructioncamera pose estimationsemantic text priorsoptical attenuationmedia-aware renderingtracking accuracy
Authors
Xiaopeng Guo, Wai Chung Tse, Yipeng Zhu, Hanwen Zhang, Huajian Huang, Sai-Kit Yeung
Abstract
Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpredictable dynamic objects. Recent feed-forward visual foundation models have demonstrated remarkable capabilities in generalized novel view synthesis and tracking. However, when directly applied to aquatic videos, optical attenuation and motion interference fatally corrupt their feature aggregation, leading to severe tracking and reconstruction failures. To overcome these limitations, we present NemoSplat, the first feed-forward 4D Gaussian Splatting framework tailored for media-aware dynamic reconstruction directly from uncalibrated marine videos. Beyond providing robust estimations of camera poses and dense scene depth, we devise a Promptable Dynamic Disentangler that utilizes a confidence-aware fusion strategy of learned dynamic probabilities and optional semantic text priors, effectively isolating massive transient entities. Furthermore, to counteract visual degradation, a Media-Aware Gaussian Predictor is formulated to jointly estimate intrinsic 3D Gaussian attributes alongside physical media parameters, rendering pristine scene appearance in a single forward pass. Additionally, we introduce a large-scale underwater dataset with massive dynamic elements to facilitate training and evaluation. Extensive experiments on our dataset demonstrate that NemoSplat achieves state-of-the-art tracking accuracy and high-fidelity rendering.