Human guided geospatial labeling speeds up drone image annotation 25 times
Human-in-the-Loop Geospatial Annotation for Rapid Dataset Construction in Field-Deployed UAV Systems
Robotics
Summary
Labeling objects in drone images by hand is very slow and hard to keep up with the large amounts of data collected. The authors developed a system called BirdsEye where human workers mark targets directly in the real world using GPS data, and then the system automatically projects these markings onto every relevant image taken from the drone. This process drastically speeds up annotation and keeps accuracy within a few centimeters. When tested on farms, this approach produced tens of thousands of labeled images much faster than manual methods, and detectors trained on this data found most of the targets in new locations with good precision.
What this means in practice
- •For agriculture monitoring teams: Generate large labeled datasets rapidly by marking plant or object locations directly in fields using GPS-equipped drones for better crop monitoring models.
- •For infrastructure inspection teams: Create annotated image data quickly by combining GPS-based real-world markings and drone imagery to improve defect or asset detection models.
Authors
Morgan Masters, Nikolaas Bender, T. Luca Altaffer, Colleen Josephson, Steve McGuire
Abstract
Real-world perception systems must adapt to changing environments, but manual image annotation cannot scale to field data volumes. We present BirdsEye, which shifts expert annotation from images to the field: an operator records target locations in world coordinates using RTK positioning and calibrated projective geometry propagates each observation to all frames where the target is visible. To quantify how well physical annotations align with image observations, we derive a first-order mapping from camera-pose uncertainty to pixel uncertainty and validate it against Monte Carlo simulation. This mapping is linear in the six per-axis pose variances, so it inverts into a sensor design tool: we give a sufficient condition converting an annotation tolerance into a convex set of admissible pose-noise budgets, a closed-form largest admissible scaling of a deployed sensor suite, and a unique per-axis pose specification under an equal-budget-share allocation. We also analyze the planar-surface approximation underlying the projection, which holds up to 10 degrees of terrain slope. By direct measurement, we show that system projection accuracy is sub-decimeter (sub-30 pixel) at AGL altitudes of 10-20m under conditions excluding sustained yawing. During an in-field case study across three agricultural sites, two field workers produced 12,524 annotated frames carrying 55,600 labels in roughly 12 hours (25.5x per-worker rate increase over manual labeling). Detectors trained on imagery collected by this workflow recovered 56-89% of in-view surveyed targets at a geographically distinct farm, at pre-registered operating points; human review of the leading configuration estimates detection precision at 83-87%, spanning three tie-break conventions for clusters carrying contradictory human verdicts.