Geometry-Driven Opti-Acoustic Co-Registration and View-Invariant Reflectivity Mapping for Side-Scan Sonar

2026-08-24Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors developed a new method to better match underwater images from sonar (acoustic) and regular cameras (optical), which is usually hard because sonar images have noise and look very different. They use 3D models of the seafloor to link the two types of images accurately and fix errors in altitude measurements. Their approach also adjusts for how sound bounces off the seabed to get true reflectivity values. This creates precisely aligned datasets without manual work, helping future research in mapping underwater habitats.

Side-Scan SonarStructure-from-Motion3D seafloor meshFirst Bottom ReturnInverse Lambertian modelReflectivity mappingCross-modal alignmentAcoustic imageryOptical imageryBenthic habitat mapping
Authors
Taqi Hamoda, Nuno Gracias
Abstract
Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlenecked by acoustic complexities such as speckle noise, shadows, and extreme viewpoint dependencies. Traditional handcrafted descriptors and modern deep learning matchers fail to bridge the physical domain gap between optical and acoustic imagery without 3D geometric constraints. To overcome these limitations, we propose a novel geometry-driven framework for pixel-level opti-acoustic co-registration and view-invariant reflectivity mapping. Our method utilizes Structure-from-Motion (SfM) to reconstruct a dense 3D seafloor mesh, acting as a geometric anchor between the visual and acoustic domains. We introduce a First Bottom Return (FBR) extraction algorithm to dynamically correct non-linear altitude drift caused by uncalibrated SfM reconstruction. Furthermore, we apply an inverse Lambertian model and a dual-Gaussian weighting function to isolate the intrinsic seabed reflectivity, effectively neutralizing slant-range propagation loss and geometric view-dependence. By deterministically associating these isolated acoustic properties with optical pixels, our pipeline generates highly accurate, strictly co-registered multi-modal datasets. This automated, physics-guided approach eliminates the need for manual annotation and paves the way for advanced self-supervised learning in benthic habitat mapping.