Video context aids navigation by linking past views to new targets

VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience

Artificial Intelligence

Summary

Navigating a space requires remembering where things are and using that knowledge to get around. The authors created a new test called VCN-Bench where a computer agent watches a video showing the start and end points before trying to find its way. This test checks if the agent can use what it saw before to understand directions and move correctly. The authors also made a method that uses both the video and real-time views to plan paths, but they found that computers still struggle to navigate even when they know the destination.

What this means in practice

  • For robotics engineers: Design robots that use prior video memories to improve indoor navigation tasks in known environments.
  • For virtual reality developers: Build VR experiences where characters navigate realistically by recalling past visual scenes combined with instructions.

Authors

Siqi Zhang, Meng Wei, Chenyang Wan, Shaohao Zhu, Shufan Shen, Xihui Liu, Zhihua Wei, Tai Wang, Jiangmiao Pang

Abstract

Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instruction following and exploration. We introduce VCN-Bench, a \textbf{V}ideo-\textbf{C}ontextualized \textbf{N}avigation benchmark for probing closed-loop spatial reasoning over prior visual experience in MLLMs. Given a prior video covering both the initial location and destination, the agent is tasked with reasoning out the instruction-specified target and navigating toward it with the inferred spatial context. Built on Matterport3D, VCN-Bench contains five instruction types, 100k training episodes, and 1,250 evaluation episodes. Navigation serves as the primary evaluation, while diagnostic goal identification helps distinguish destination-resolution errors from subsequent navigation failures. We further propose MV-DualVLN, a planning-oriented baseline that jointly leverages prior video and in-episode observations. Experiments reveal limited navigation performance, a substantial destination-resolution-to-navigation gap, and frequent navigation failures even after correct destination identification.