CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments

Robotics

Summary

The authors created a system that helps underwater robots find their way through tricky caves without bumping into walls. Instead of relying on clear pictures which are hard to get underwater, their system uses smart reasoning with data from different sources like light, shapes of passages, and sonar measurements. This helps the robot understand which way to go safely. Tests in realistic simulations showed their method successfully navigates caves while avoiding obstacles.

Authors

Zhenqi Wu, Yuanjie Lu, Yisheng Zhang, Miao Yu, Xuesu Xiao, Jaejeong Shin, Xiaomin Lin

Abstract

Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features for localization and mapping. In underwater caves, however, visual degradation can undermine feature-based localization, sonar-based mapping may yield overly conservative obstacle representations, and communication constraints preclude real-time human guidance. To address these limitations, we propose an autonomous underwater cave navigation framework that leverages a vision-language model (VLM) with Chain-of-Thought (CoT) reasoning to infer navigable directions from environmental cues, including light intensity gradients, passage morphology, and geometric complexity, captured through multimodal inputs comprising RGB imagery, depth maps, and sonar-based vertical-clearance measurements, thereby supporting safe 3D navigation through confined cave passages. High-fidelity simulations across multiple cave topologies demonstrate that the proposed framework completes all evaluated end-to-end traversals without collisions while maintaining safe clearance from cave boundaries.