Spheriverse improves 3D scene understanding from spherical images
Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild
Computer Vision and Pattern RecognitionRobotics
Summary
Understanding 3D scenes using spherical images is challenging because these images wrap around in angles while real-world scenes use regular coordinates. The authors created a large dataset called Spheriverse with pairs of spherical images and 3D LiDAR scans to help solve this problem. They also built SphereOcc, a new method that better connects spherical image information to 3D space, improving accuracy in predicting what's where in a scene. Their approach outperforms earlier methods in various tasks like identifying objects and mapping scenes under different conditions.
spherical images3D scene understandingLiDARsemantic occupancy predictionCartesian coordinatesangular domainvoxel featuressemantic mapping3D object detectionrange-azimuth geometry
Authors
Fei Teng, Sheng Wu, Mengfei Duan, Guoqiang Zhao, Junhui Ma, Kai Luo, Siyu Li, Hao Shi, Zhiyong Li, Kailun Yang
Abstract
Spherical observations provide global visual context for 3D scene understanding. However, visual information is encoded in an angular domain, whereas the physical world is represented in Cartesian coordinates. This cross-space representation gap complicates geometric correspondence and semantic evidence aggregation. To delve into this challenge, we introduce Spheriverse, comprising $64,400$ temporally aligned spherical image-LiDAR pairs organized into 644 sequences. The dataset spans diverse scenes, illumination, and weather conditions, with fine-grained semantic classes. We further establish benchmarks for semantic occupancy prediction, semantic mapping, and 3D object detection, evaluating 30+ methods through overall and scene-wise comparisons. For dense prediction, we propose SphereOcc, an occupancy framework that couples spherical geometry modeling with semantic evidence retrieval. Cartesian-Spherical Representation Remodeling (CSRR) incorporates spherical range-azimuth geometry into Cartesian voxel features through region-wise modulation. Spherical Evidence Re-querying (SER) then conditions queries on voxel content and range-height-azimuth geometry to adaptively retrieve relevant semantic evidence from source spherical image features. SphereOcc achieves 13.91% mIoU and 24.65% GeoIoU, outperforming the respective best-performing methods, TPVFormer and SurroundOcc, by 1.70 and 2.10 percentage points. It also ranks first in both metrics across all five scenes, with consistent advantages across the evaluated spatial partitions and reduced fields of view. The established benchmark and source code will be available at https://feit-feiteng.github.io/Spheriverse.