Summary
3D reconstruction from photos often assumes all images come from the same type of camera, which isn't true in real life where different lenses like fisheye or 360-degree cameras might be used together. This paper introduces MEOW, a system that can take a group of photos from various camera types all at once and build a 3D model without needing information about the cameras beforehand. The key idea is to train the system on a wide range of synthetic images that simulate different cameras, enabling it to work on real-world photos without extra setup. As a result, MEOW can reconstruct scenes more flexibly and accurately than previous methods that required camera labels or handled fewer images at a time.
What this means in practice
- •For augmented reality developers: Create accurate 3D models from mixed camera inputs without manual camera calibration for rich AR experiences.
- •For robotics teams: Enable robots to build 3D maps from heterogeneous camera setups in real environments rapidly and without extra sensor input.
Authors
Qiaoge Li, Yifan Zhan, Haijun Yang, Haiyang Liu, Yiyi Cai, Chenchi Luo
Abstract
Real-world capture is heterogeneous: perspective, fisheye, and $360^\circ$ panoramic images can coexist within a single reconstruction task, yet most feed-forward 3D reconstruction models assume perspective imagery and a uniform input representation. Recent models handling several camera types are either informed of the camera type for each view or reconstruct one image pair at a time. No single-pass method reconstructs mixed-camera tuples containing full panoramas from images alone. We present MEOW, a feed-forward system that jointly reconstructs metric pointmaps and camera poses from one N-view tuple mixing perspective, fisheye and full-panorama images, in a single forward pass from images alone: no calibration, distortion parameters, camera-type labels or poses are supplied for any view. Our guiding design philosophy is to treat heterogeneous-camera reconstruction as a data-adaptation problem rather than an architectural redesign. MEOW retains a perspective-pretrained backbone and learns heterogeneous cameras entirely from a procedural data engine, which renders each scene across a continuous manifold of camera models with exact rays and depth, and certifies covisibility for every camera-sampled training tuple. Trained on synthetic tuples only, MEOW transfers zero-shot to real captures: on heterogeneous 2D3DS tuples it achieves 79.9 mAA@30 against 53.8 for Wid3R given the camera type of every view; on our laser-scanned mixed-camera benchmark it registers every four-view mixed tuple with 79.4 AUC@30. The data engine, benchmark, and complete evaluation pipeline will be released.