ExcavaTwin maps terrain and objects for autonomous excavation tasks
ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation
Robotics
Summary
Accurately knowing the shape and type of ground is crucial for robots digging in the dirt. The authors created ExcavaTwin, which uses multiple camera images to build detailed 3D maps showing both the terrain shapes and what kind of surfaces or objects are present. Their method works without needing special training for digging tasks because it uses existing vision tools and geometry rules to combine information reliably. Tests in real and public excavation scenes showed it provides fairly accurate and timely updates about changing ground conditions.
What this means in practice
- •For construction equipment operators: Use ExcavaTwin to obtain up-to-date 3D maps combining terrain shape and identifying ground features for safer and more precise autonomous digging operations.
- •For robotic navigation teams: Implement geometry-informed semantic mapping from RGB cameras to enhance terrain understanding for outdoor robotic tasks involving changing environments.
Authors
Yu Deng, Lingshan Zeng, Tong Hu, Rushi Dai
Abstract
Autonomous excavation requires a spatial representation that jointly captures terrain geometry and task-relevant semantics. Existing excavation mapping is largely elevation-centric, while generic semantic models remain unstable in unstructured outdoor scenes. We present ExcavaTwin, a pure-vision geometry-guided semantic elevation mapping framework without excavation-specific training. Given multi-view RGB images, the framework: 1) reconstructs scene geometry and semantic observations using frozen vision models; 2) derives terrain and non-terrain geometric support; 3) performs geometry-constrained multi-view semantic fusion to suppress implausible predictions and recover incomplete observations; and 4) projects the fused state into a task-oriented semantic elevation map. Experiments on public datasets and real excavation scenes demonstrate reliable geometric and semantic perception. In real excavation, the system achieved an average update interval of approximately 1.4 s and a mean elevation error of 12.74cm in dynamically modified regions. Larger errors mainly occur during rapid terrain changes and transient visual disturbances caused by machine motion.