Papers for

drone navigation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Visual place recognition improves with reliability guided aggregation

Calibrating Retrieval Geometry: Reliability-Guided Training-Free Aggregation for Visual Place Recognition

Abstract: Frozen visual foundation models provide transferable features for visual place recognition, but fixed aggregation can suppress useful distinctions in new environments. We introduce TFA, a reliability-guided, training-free aggregation method requiring neither place labels nor task-specific weight updates. Our key observation is that reproducible retrieval need not be discriminative: independent codebooks can consistently retrieve a few database hubs. TFA combines cross-codebook agreement, retrieval coverage, and spectral statistics to control residual assignment, spectral shaping, and global-feature fusion. Its spectral kernel exactly recovers original descriptor similarity at zero intervention. Database-only TFA fixes its rules before accessing queries; TFA-C64 uses 64 disjoint unlabeled target images to calibrate retrieval for subsequent queries. Across 20 ground protocols with a fixed DINOv2-B backbone and matched resolution, database-only TFA improves Recall@1 over AnyLoc by 17.39 percentage points on MSLS-val and 9.55 on SPED. C64 mitigates failures of database-only calibration in driving environments. Across eight aerial/cross-view protocols, TFA achieves the highest Recall@1 among compared training-free heads in 14 of 16 DINOv2/DINOv3 backbone-protocol combinations. In a separate native-system comparison, DINOv2-G-based TFA-C64 reaches 91.46% Recall@1 on Pitts30k and 76.29% on VPAIR, outperforming the displayed training-free comparators on all five benchmarks. These results show that reliability-guided aggregation can recover additional retrieval capability from frozen representations, providing a practical baseline for new environments with scarce place supervision.

Tue 22 SeptComputer Vision and Pattern Recognition
The gist
Visual place recognition helps computers identify where a photo was taken by comparing image features. The authors found that fixed methods for combining these features can miss important details in new environments. They created a method called TFA that groups features reliably without needing extra training or labels. This method adjusts based on how consistent different feature sets are and can improve recognition accuracy in various challenging settings.
Open → 2609.25937v1

UAVs use predicted views to explore hidden spaces faster and better

WOLF: World Model Guided LiDAR Exploration with Predictive Frontiers

Abstract: LiDAR-based unmanned aerial vehicle (UAV) exploration builds maps by continually selecting where to observe next. However, decisions based on the measured map provide limited foresight into spatial continuations behind occlusions, leaving potentially informative directions unrecognized. We present WOLF, a world-model-guided framework that predicts future observations to enhance autonomous exploration. In the training stage, a recurrent world model learns observation dynamics from exploration trajectories, with recurrent memory retaining the spatial context needed to interpret partial observations across successive views. Building on this context, the model combines observation history with candidate motions during exploration to predict local occupancy and visibility. To guide further sensing, a predictive frontier generation mechanism then aligns and fuses these predictions using confidence, branch agreement, and observation quality to identify promising regions. The resulting predictive frontiers join measured ones to guide geometric viewpoint selection and trajectory generation, while new scans update subsequent predictions. In simulations, our method reduces mean terminal time by 10.9% relative to EPIC in Garage at comparable coverage and increases mean coverage from 42.12% to 98.35% in Tunnel. Real-world experiments further demonstrate onboard deployment of the learned model for online inference during physical flight.

Sun 20 SeptRobotics
The gist
Mapping unknown places with drones using laser sensors can miss important spots hidden behind obstacles. The authors present WOLF, a method where a smart model predicts what is likely to be behind these obstacles based on past observations, helping the drone decide where to look next. This prediction guides the drone more effectively, speeding up exploration and covering more area. They tested WOLF in simulations and real flights, showing improved speed and completeness compared to a prior method.
Open → 2609.23656v1

PerSeM improves long-term semantic mapping for UAVs with persistent memory

PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping

Abstract: Open-vocabulary segmentation enables rich semantic perception for UAVs, but frame-wise predictions can remain temporally inconsistent across repeated observations and changing viewpoints. We present PerSeM, a training-free persistent semantic memory framework for long-horizon open-vocabulary UAV mapping. PerSeM associates frame-wise semantic observations with persistent world-space voxels and constructs a majority-based semantic memory, which is conservatively refined through history-preserving spatial refinement, trust-aware replay, and context-guided verification. Experiments on the Forest and UAVScenes benchmarks show that persistent 3D memory provides substantial gains in semantic correctness and temporal stability over frame-wise predictions. Beyond this strong persistent-memory baseline, PerSeM provides consistent additional improvements, improving both semantic accuracy and temporal stability across all five evaluated UAVScenes sequences. Analysis using regions identified independently of the final PerSeM predictions further shows that these gains are concentrated in semantically difficult and temporally unstable regions, where majority-based memory is most likely to remain uncertain. These results demonstrate that persistent 3D aggregation provides a strong foundation for long-horizon semantic mapping, while conservative refinement of uncertain memory states can provide additional improvements without retraining or additional neural-network inference.

Thu 17 SeptComputer Vision and Pattern RecognitionRobotics
The gist
UAVs can see and label things around them but often get inconsistent results when looking repeatedly from different angles. The authors developed PerSeM, which keeps a memory of what the UAV has seen over time, combining many observations to get more accurate and stable labels. This method doesn’t need to be trained again and improves the UAV’s understanding especially in tricky places. Tests showed PerSeM makes the UAV’s semantic maps both more accurate and reliable over long periods.
Open → 2609.19542v1

Uncertainty-aware AI improves off-road robot navigation routes

UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner

Abstract: Vision-language models (VLMs) can generate routes directly from aerial imagery for off-road navigation, but their predictions provide no indication of reliability. We present UDAV, an Uncertainty-Driven Adaptive VLM Waypoint Planner for UAV-guided UGV navigation. UDAV draws multiple stochastic trajectory predictions, selects their medoid as a self-consistent nominal route, and estimates predictive uncertainty from their spatial dispersion. When the maximum uncertainty across interior waypoints exceeds a threshold, UDAV invokes a reconsideration stage; otherwise, it returns the medoid directly. We evaluate UDAV on 400 held-out trajectory queries from two UAV flights. Stochastic medoid selection reduces the mean average displacement error (ADE) from 147.4 pixels for a deterministic prediction to 115.9 pixels. The complete planner achieves a mean ADE of 110.4 pixels, a 25.1% reduction relative to deterministic planning, while producing valid trajectories for all queries. UDAV also yields the lowest 90th- and 95th-percentile errors among all evaluated configurations, including a higher-budget K=10 consensus baseline. Relative to the K=5 medoid, UDAV reduces these errors from 225.3 and 326.0 pixels to 199.0 and 290.8 pixels, respectively. These results demonstrate that stochastic VLM predictions provide both a stronger nominal route and an actionable uncertainty signal for selectively mitigating large planning errors.

Mon 14 SeptRoboticsArtificial Intelligence
The gist
Navigation systems for robots using aerial images can get confused and make mistakes without knowing how confident they are. The authors created a method called UDAV that makes many route guesses, finds the most typical one, and measures how reliable it is. If the route looks uncertain, UDAV rethinks the plan to avoid big errors. Tests show this idea makes navigation more accurate and safer for off-road vehicles guided by drones.
Open → 2609.16368v1