Quantized vision language model runs efficiently on edge devices for navigation

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

RoboticsComputer Vision and Pattern Recognition

Summary

Robots that navigate using language and vision usually need powerful computers to work well. The authors created EdgeVLN, a smaller, faster version that fits on limited hardware like robot edge devices. It uses clever tricks like reducing number size and a lightweight stopping predictor to save memory and energy while keeping good navigation success. Testing showed it works almost as accurately as the big model but runs much faster and uses less power on a common robot computer.

What this means in practice

  • For robot developers: Deploy vision-language navigation models on low-power robots with real-time stopping accuracy and low memory use.
  • For embedded systems engineers: Implement efficient inference of multimodal navigation models on memory- and energy-constrained edge hardware.

Authors

Rithvik Jonna, Man Namgung, Aakash Gurram, Tinoosh Mohsenin

Abstract

Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavior. We introduce EdgeVLN, a runtime-aware, deployment-ready quantized VLN model that closes this gap. EdgeVLN combines a quantized StreamVLN model with Latent Trajectory Termination Extractor (LATTE), a lightweight causal transformer that improves real-time stopping by predicting a Stop Action verifier rank. Both execute through our llama.cpp VLN driver, which reconstructs streaming context and prunes memory tokens on-board. We characterize a pretrained StreamVLN backbone across weight quantization from 8 to 2 bits and multiple inference runtimes to identify a feasible operating point. LATTE reuses backbone hidden states within the budget freed by quantization, requiring neither a second vision encoder nor an additional backbone forward pass. We evaluate six backbone precisions and seven candidate stop heads on BF16 and IQ4 NL across all 1,839 R2R VLN-CE val-unseen episodes. We measure success rate (SR) in simulation and latency, energy, and resident memory on an NVIDIA Jetson Orin NX 16 GB. LATTE achieves our highest SR, 58.02 percent on the deployed 4-bit model, exceeding the BF16 baseline with only 0.013 s additional latency per navigation step. Four-bit formats achieve nearly identical SR, but step energy varies 36.8 times by execution path. Only IQ4 NL under our VLN driver fits the board, using 11.35 GB resident memory while running 20.8 times faster and using 13.3 times less energy than storage-streamed BF16. INT2 collapses. Runtime selection, memory-token pruning, and quantization are essential for efficient edge deployment.