SparseNav improves indoor robot navigation with selective map labeling

SparseNav: Instruction-conditioned Sparse Semantic Perception for Training-Free Vision-Language Navigation

Robotics

Summary

Indoor robots need to understand both language instructions and their surroundings to find their way. The authors created SparseNav, a system that only adds important landmarks to a map based on the current instruction, saving computing power and avoiding confusion. It works without extra training and was tested successfully in simulated environments and on a real robot navigating indoors. This selective approach helps robots follow complex directions more efficiently.

What this means in practice

  • For robotics engineers: Enable robots to follow spoken indoor navigation instructions using a lightweight map updated only with relevant landmarks as needed.
  • For mobile robot developers: Implement training-free vision-language navigation on quadruped robots for efficient real-time landmark grounding without prebuilt maps.

Authors

Quanhua Chen, Juhan Kang, Runfeng Lin, ZiFei Zhang, Enquang Feng, Chunran Zheng, Xiwang Dong, Jiarong Lin

Abstract

Map-based vision-language navigation (VLN) relies on persistent spatial representations to connect language understanding with geometric planning. However, acquiring semantics beyond the needs of the current instruction can introduce unnecessary perception cost and irrelevant annotations. Continuously accumulating unrelated objects may not only waste computation, but also clutter the visual-spatial representation consumed by the vision-language model (VLM) planner. To address this problem, we present SparseNav, a training-free framework that follows a less-is-more principle for semantic navigation. SparseNav persistently maintains a lightweight geometric bird's-eye-view (BEV) map and sparse landmark memory, acquiring new semantics on demand using the active sub-instruction to decide what is worth grounding. An instruction manager first tracks navigation progress and identifies the active landmark query. An instruction-conditioned perception mechanism then invokes open-vocabulary segmentation when the queried landmark is visible and its metric location can inform the next decision. The resulting landmark memory supports VLM selection among hybrid frontier and local directional waypoint candidates. Without any additional training, SparseNav achieves success rates of 42.8% on R2R-CE and 40.7% on RxR-CE, both on the Val-Unseen splits. Controlled ablations examine semantic perception strategies and the contributions of individual framework components. Furthermore, we successfully deployed SparseNav on a Unitree Go2 quadruped equipped with an Intel RealSense D455 RGB-D camera for geometric mapping and landmark grounding and a Livox MID-360 LiDAR for localization, without a prebuilt map. We validated its effectiveness across multiple indoor environments using instruction-conditioned waypoint navigation.