Multi view stereo achieves better 3D models with sequence to sequence approach

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

Computer Vision and Pattern Recognition

Summary

Figuring out the shapes of objects from pictures taken from different angles is tricky and often leads to errors. The authors propose a new way to solve this by treating the problem like translating sequences, predicting 3D details for all pictures together instead of just one. They use a special network that understands camera settings to improve accuracy across all views at once. Tests show that their method makes better 3D reconstructions than previous techniques.

What this means in practice

  • For 3d scanning engineers: Create more accurate 3D models from multiple images by jointly processing all camera views with better geometric consistency.
  • For robot perception developers: Enhance scene understanding in robots by improving depth and geometry estimation from multiple camera angles.

Authors

Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua

Abstract

Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied. In this paper, we study the multi-view stereo (MVS) problem with known camera parameters and propose a novel approach that bridges conventional MVS and FF methods. Rather than casting MVS as a sequence-to-one mapping that predicts depth only for a single reference view, we reformulate it as a sequence-to-sequence task, akin to FF models, that jointly predicts geometry for all input views. We introduce a global transformer-based architecture with two components that explicitly exploit camera-induced priors: ray-map embeddings that inject camera parameters into image patch tokens, making the transformer camera-aware, and a unified global cost volume that replaces conventional per-view cost volumes to jointly capture 3D structure across all views. Extensive experiments on multiple public benchmarks show our approach achieves state-of-the-art performance, surpassing both MVS and FF reconstruction baselines.