Papers for

infrastructure inspection teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Human guided geospatial labeling speeds up drone image annotation 25 times

Human-in-the-Loop Geospatial Annotation for Rapid Dataset Construction in Field-Deployed UAV Systems

Abstract: Real-world perception systems must adapt to changing environments, but manual image annotation cannot scale to field data volumes. We present BirdsEye, which shifts expert annotation from images to the field: an operator records target locations in world coordinates using RTK positioning and calibrated projective geometry propagates each observation to all frames where the target is visible. To quantify how well physical annotations align with image observations, we derive a first-order mapping from camera-pose uncertainty to pixel uncertainty and validate it against Monte Carlo simulation. This mapping is linear in the six per-axis pose variances, so it inverts into a sensor design tool: we give a sufficient condition converting an annotation tolerance into a convex set of admissible pose-noise budgets, a closed-form largest admissible scaling of a deployed sensor suite, and a unique per-axis pose specification under an equal-budget-share allocation. We also analyze the planar-surface approximation underlying the projection, which holds up to 10 degrees of terrain slope. By direct measurement, we show that system projection accuracy is sub-decimeter (sub-30 pixel) at AGL altitudes of 10-20m under conditions excluding sustained yawing. During an in-field case study across three agricultural sites, two field workers produced 12,524 annotated frames carrying 55,600 labels in roughly 12 hours (25.5x per-worker rate increase over manual labeling). Detectors trained on imagery collected by this workflow recovered 56-89% of in-view surveyed targets at a geographically distinct farm, at pre-registered operating points; human review of the leading configuration estimates detection precision at 83-87%, spanning three tie-break conventions for clusters carrying contradictory human verdicts.

Wed 23 SeptRobotics
The gist
Labeling objects in drone images by hand is very slow and hard to keep up with the large amounts of data collected. The authors developed a system called BirdsEye where human workers mark targets directly in the real world using GPS data, and then the system automatically projects these markings onto every relevant image taken from the drone. This process drastically speeds up annotation and keeps accuracy within a few centimeters. When tested on farms, this approach produced tens of thousands of labeled images much faster than manual methods, and detectors trained on this data found most of the targets in new locations with good precision.
Open → 2609.28767v1

AirSplan improves safe quadrotor flight in complex 3D environments

AirSplan: Risk-Aware Motion Planning for Quadrotors in Cluttered 3D Gaussian Splats

Abstract: Quadrotors are increasingly deployed in applications such as agriculture, infrastructure inspection, and maintenance. In each of these applications, the robot must navigate complex scene geometry while remaining strictly collision-free. Unlike in ground domains, even minor collisions for aerial vehicles can result in the loss of the robot. This safety requirement induces a pair of technical challenges. First, the environment must be represented with sufficient fidelity to encode complex structure, even when no ground-truth obstacle data is available. Second, a motion planner must leverage this representation to determine a collision-free path to the goal. This paper proposes a system that addresses these complementary challenges. The proposed method, AirSplan, adopts a normalized variant of 3D Gaussian Splatting that encodes high-fidelity scene geometry. It then applies a novel reachability-based motion planner that leverages the differential flatness of quadrotors to compute continuous-time collision constraints that tightly overapproximate the robot's occupancy. Experiments demonstrate that AirSplan successfully finds a path in 81.2% of challenging test cases, a significant improvement over the nearest baseline method's 51.2%.

Fri 18 SeptRobotics
The gist
Flying small drones near obstacles is risky because even tiny collisions can cause crashes. The authors created AirSplan, a system that builds detailed 3D maps from uncertain data and then plans safe flight paths to avoid collisions. Their method uses a special way to represent the environment and optimizes the drone's path considering its flying dynamics. Tests show AirSplan finds collision-free routes much more often than previous methods.
Open → 2609.21226v1

LLM enhances tunnel lining inspection accuracy from 3D point clouds

R4Tun: LLM-guided adaptive segmental tunnel lining segmentation in point clouds

Abstract: Automated inspection of segmental tunnel linings requires adaptive segmentation from 3D point clouds, yet expert-tuned pipelines often degrade when tunnel conditions vary. This paper presents R4Tun, a large language model (LLM)-driven adaptation framework that extends an expert-designed pipeline (SAM4Tun) with bounded parameter tuning informed by structured context: memory ($m$), state ($s$), and knowledge ($k$). Evaluated on 30 selected Seg2Tunnel subsets (13 regular, 17 complex) across three LLMs, the full $m+s+k$ design raised mean Intersection-over-Union (mIoU) from 0.18 to 0.43--0.48 and overall accuracy (OA) from 0.42 to 0.59--0.65 relative to the static SAM4Tun baseline, with the near-reference regular (staggered) subsets reaching mIoU 0.784--0.796 across LLMs. Across 270 (30 tunnels $\times$ 3 different LLMs $\times$ 3 context settings) runs, the LLMs showed similar parameter-adjustment trends (with overlapping 95\% CIs on mean gains) and consistently adjusted a shared set of critical parameters. These results support R4Tun as a controlled, label-free, cross-LLM adaptation mechanism in the tested SAM4Tun--Seg2Tunnel setting, demonstrating consistent accuracy gains; we position R4Tun as a mechanism contribution rather than a deployable final-inspection system, in which each bounded parameter change is auditable via logged rationales.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Inspecting tunnels is important but tricky because conditions change and automatic methods often struggle. The authors developed R4Tun, a system that uses large language models (LLMs) to adjust the inspection process automatically based on the tunnel's situation. This makes the system much better at identifying parts of the tunnel lining without needing extra labeled data. Their tests showed clear improvements in accuracy across different tunnel types and LLM types.
Open → 2609.11360v1