Point diffusion mamba improves 3d reconstruction with less data
Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity
Computer Vision and Pattern Recognition
Summary
Reconstructing 3D shapes from a single 2D image is very hard, especially when there isn’t much training data available. The authors developed Point Diffusion Mamba (PDM), a new method that combines two techniques to better predict 3D shapes even with limited data. PDM uses a clever way to understand both the overall shape and small details of 3D objects, and it improves accuracy by mixing different features and sampling strategies. Tests show PDM works better than previous methods on standard 3D shape datasets.
What this means in practice
- •For 3d graphics developers: Generate accurate 3D object models from single 2D images even when training data is limited.
- •For augmented reality app teams: Improve object reconstruction quality for AR experiences on devices with limited dataset availability.
Authors
Wei Zhou, Xinzhe Shi, Xingxing Hao, Xing Hao, Kang Li, Jinye Peng, Ying He
Abstract
While single-view 3D reconstruction has seen significant progress, extrapolating complex 3D structures from inherently ambiguous 2D observations remains fundamentally ill-posed, particularly in the critically underexplored data-scarce regime. To address this challenge, we propose Point Diffusion Mamba (PDM), a method that integrates the generative power of diffusion models with the efficiency of state-space model for single-view 3D reconstruction under data-scarce conditions. Specifically, PDM employs a lightweight reconstruction module tailored to handle unordered point-cloud inputs effectively. By combining a Local Geometric Aggregation module with Mamba blocks, our approach jointly models global geometric structures and local details. In 3D reconstruction, each point in the initial noisy input requires a precise prediction, yet the high-level features extracted by the Mamba module capture only abstract semantic information from sparse points. To bridge this gap, we introduce the Hierarchical Feature Integration Network, which fuses high-level semantic and local geometric features for each point, overcoming the limitations of token-based point-cloud reconstruction. Furthermore, we propose a Dynamic Weighted Sampling strategy that adaptively unifies 3D generation with single-view reconstruction by leveraging generative priors to enhance reconstruction quality. Experimental results on the ShapeNet and Pix3D benchmarks demonstrate that PDM outperforms state-of-the-art methods, providing an effective solution for 3D reconstruction under data-scarce settings. Code is available at: https://github.com/NWUzhouwei/PDM.