Omni-language models enable zero-shot audio visual navigation

RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation

Artificial Intelligence

Summary

Navigating by combining what you hear and see is hard for robots, especially without training on specific tasks. The authors developed RAO-Nav, a method that uses large language models capable of understanding both sounds and visuals to help robots reason and move around new places without prior training. They also created a way to improve the robot's navigation decisions by encouraging it to focus on the most relevant information. Their approach works better than specialized models on existing tests without any extra training. They also suggest a new challenge to test how well these models can handle broader navigation instructions.

What this means in practice

  • For robotics developers: Build audio-visual robots that navigate new environments without task-specific training data by using omni-language models and latent navigation reasoning.
  • For smart home device makers: Enhance smart assistants to locate objects or navigate homes by integrating zero-shot audio-visual reasoning capabilities.$Commercial implications: Enables consumer devices to perform general audio-visual navigation tasks without costly retraining, improving user experience and functionality.

Authors

Qilang Ye, Meng Liu, Yu Zhou

Abstract

We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: https://github.com/rikeilong/OmniAV\_Nav.