Satellite images enable city scale drone navigation benchmark

SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery

Computer Vision and Pattern RecognitionRobotics

Summary

Navigating drones across large cities is hard because it needs memory and understanding of maps. The authors created SatNav, a big set of city journeys using satellite images to simulate what drones see from above. They tested different computer systems on SatNav and found city-wide navigation still difficult. They also made a tool called SwiftVLN to try out different memory methods. Models trained on these satellite images can also work on real drone flights, showing this approach can be useful in practice.

What this means in practice

  • For drone software developers: Develop drone navigation systems that understand city-scale routes using satellite-image-based training data for better long-range performance.
  • For mapping and surveying teams: Use satellite-based benchmarks to improve models that guide UAVs for extensive urban data collection and inspection tasks.

Authors

Jiajun Jiang, Chunliang Hua, Zichun Chen, Yanxing Wu, Zeyuan Yang, Jie Song, Xiao Hu

Abstract

Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, limiting geographic diversity and episode scale. To address this, we introduce SatNav, a scalable, long-horizon UAV VLN benchmark built from high-resolution satellite imagery. SatNav targets city-level navigation missions and uses satellite crops as approximations of UAV nadir views for visual observations. Through an automated cue-to-episode pipeline, SatNav constructs 118K episodes from 59 scenes across 18 cities, with an average trajectory length of 379 m. To stress-test long-horizon memory and geospatial reasoning, SatNav defines three task families: Boundary, Landmark, and Route, targeting loop progress tracking, landmark-based spatial grounding, and route following with counting cues. Benchmarking classical VLN agents and recent agents based on large vision-language models (LVLMs) on SatNav shows that city-scale navigation remains challenging. We further introduce SwiftVLN, a modular framework with switchable memory components, and conduct systematic memory-design ablations. Finally, satellite-to-UAV transfer experiments show that satellite-trained navigation models can operate on real-flight UAV observations, showing the practical relevance of SatNav. Our project page: https://eku127.github.io/SatNav/