Underwater robots map pools accurately using sonar and odometry

Odometry-Aided Real-Time Mapping for Underwater Robots Using Forward-Looking Sonar

Robotics

Summary

It's hard for underwater robots to see clearly because water messes up light, so instead they use sonar, which sends out sound waves to create images. The authors created a method to clean up these sonar images and connect shapes to understand the surroundings better. They also use data from the robot’s movements to build a map of underwater spaces in real time. Their method works well in a pool, creating maps that are accurate to within a few centimeters.

What this means in practice

Authors

Siyuan Du, Kanzhong Yao, Youdong Wang, Yingqi Liu, Qingwen Liu, Qunhui Yang, Zhe Sun, Xuelong Li

Abstract

Reliable perception is essential for underwater vehicles operating in complex environments, where light attenuation and scattering often degrade visibility and compromise optical sensing. Forward-looking sonar (FLS) offers an alternative by providing high-frame-rate acoustic imaging under poor optical conditions. However, real-time FLS mapping remains challenging due to unresolved target elevation, spatially non-uniform noise, and fragmented target boundaries, which hinder feature extraction and introduce geometric ambiguity during projection. To address these challenges, we propose a cascaded feature reconstruction pipeline combining fast Fourier transform (FFT)-based denoising, fast multiscale constant false alarm rate (MCFAR) detection, and gradient-adaptive boundary connection to extract geometric features from degraded sonar images with low latency. We integrate attitude-aware geometric projection with incremental occupancy accumulation to construct a depth-referenced 2.5D map for local mapping in confined underwater environments. The sonar's vertical position is referenced to an external sensor, while target elevation is assigned under an explicit geometric assumption rather than measured directly by FLS. Experiments in a 3 m X 5 m pool demonstrate centimeter-scale planar mapping accuracy, with a root-mean-square error (RMSE) below 3 cm across three sequences and an average processing time of 42.4 ms per frame.