NavJev speeds up visual navigation with compact action-focused memory
NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory
Robotics
Summary
Vision-and-Language Navigation systems help robots find their way by combining pictures and instructions. The authors found that existing methods take too long because they rethink everything at every step. They created NavJev, which quickly summarizes the important visual details tied to possible moves and remembers useful features from earlier steps to decide faster. This approach makes navigation much quicker without sacrificing too much accuracy.
What this means in practice
- •For robotics engineers: Build indoor navigation robots that respond faster by using compact visual summaries to choose actions efficiently.
- •For mobile app developers: Create augmented reality apps that quickly interpret instructions and surroundings to guide users with minimal delay.
Authors
Kai Sheng, Liuyi Wang, Jinlong Li, Haojie Dai, Chengju Liu, Qijun Chen
Abstract
Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Specifically, Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods. The project page is available at https://kai-sheng-caesar.github.io/NavJev/.