Model improves 3d object separation from sparse multiple images

SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views

Computer Vision and Pattern Recognition

Summary

Separating objects from 3D scenes is challenging, especially when working with a few pictures taken from different angles. This paper presents a new method that uses object masks from a 2D segmentation model to guide 3D reconstruction, helping to focus on one object at a time. Their approach improves the accuracy of 3D shapes and makes it easier to distinguish individual objects in a scene. They also developed a lightweight network to pick the best viewpoint for reconstruction, leading to more stable and reliable results.

What this means in practice

Authors

Langxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao, Tianhan Gao

Abstract

With the rising demand to decouple objects from 3D scenes, we propose SAMV-DUSt3R, an end-to-end model that injects SAM2 2D masks into MV-DUSt3R reconstruction. A Cross Flow Mask Block uses these masks to steer the network toward the target instance, jointly improving shape accuracy and achieving object-level disentanglement without multi-stage pipelines. To ensure reconstruction stability, a lightweight Spatial RankGNN selects the optimal reference view with a selection accuracy of 73.5\%. Extensive experiments demonstrate that our method boosts average reconstruction precision by 11\% across various metrics compared to state-of-the-art baselines. These results reveal a strong instance-disentanglement capability and clear benefits for driving, robotics, AR/VR, and heritage digitisation.