GPT-6-Astra shows strengths and limits in zero-shot robot navigation
GPT-6-Astra in a Navigation Workflow: Behavioral Analysis in Zero-Shot Vision-and-Language Navigation in Continuous Environments
Robotics
Summary
This study looks at how GPT-6-Astra can understand spoken or written directions and move through a space it has never seen before, without prior training for navigation. The authors found that the system can describe landmarks and connect past moves to instructions, sometimes asking for better views before deciding what to do. However, the system struggles to stop at the right time or keep making progress after recognizing problems, showing a gap between understanding the task and completing it. This means GPT-6-Astra can make smart local decisions but has trouble finishing navigation tasks fully on its own.
What this means in practice
- •For robot developers: Improve natural language instruction handling by robots navigating unknown environments without customized training.
- •For smart home device makers: Integrate vision-and-language models to enable devices to follow spoken navigation commands more reliably indoors.
Authors
Guangzhao Dai, Qi Wu, Bin Zhu
Abstract
We study GPT-6-Astra in a zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) system, where it interprets instructions, assesses its surroundings, and proposes actions. The system uses a common observation--decision--execution workflow with direct model API calls, without a packaged agent harness or navigation-specific fine-tuning. In this workflow, each request receives selected observations, execution feedback, and retained progress records. Evaluation covers the complete system, including context management and action control. We evaluate the system on 50 of the 100 R2R-CE val-unseen episodes used by Open-Nav. It achieves a success rate of 52.0\%, an SPL of 48.9\%, and an nDTW of 70.8\%. Our analysis highlights three findings. First, recorded responses link landmarks and earlier actions to instructions using observations and supplied history. Second, reviews include requests for additional views and revisions of uncertain judgments. Third, the results suggest a gap between task understanding and autonomous completion: an unfinished crossing is recognized while rotation continues. At termination, 36.0\% of episodes succeed with a workflow-accepted STOP, while another 16.0\% meet the distance criterion at the step limit. These results highlight a central challenge: translating correct local judgments into sustained progress and appropriate stopping.