Legged robot improves object coverage using pose and omnidirectional sensing

Pose-aware Legged Robot Semantic Exploration with Omnidirectional Perception in Confined Unknown Environments

Robotics

Summary

Exploring tight, unfamiliar spaces is hard for robots because their cameras and sensors often miss parts of objects, especially the tops. The authors created POSE, a system that helps a legged robot tilt and change its posture smartly to see objects better using all-around cameras and lasers. This system decides how the robot should position itself to get the most useful views without wasting time. Tests in simulations and a real machine shop showed POSE finds more object surfaces faster than older methods. This could help robots inspect places people find hard to reach.

What this means in practice

  • For industrial inspection teams: Improve inspection of complex machinery in tight spaces by enabling legged robots to adjust posture for better object surface coverage with omnidirectional sensors.
  • For search and rescue operators: Enhance search missions in confined or collapsed environments by using pose-aware robot movements to comprehensively scan obstacles and trapped objects.

Authors

Xiaoyang Zhan, Shiyu Chen, Kenji Shimada

Abstract

Semantic exploration in confined environments requires both environment mapping and detailed observation of target objects. For ground robots, limited sensor vertical fields of view and restricted standoff distances can leave upper object surfaces unobserved from planar viewpoints. Body tilting can improve coverage, but additional observations and posture transitions increase mission time. To address this trade-off, we present POSE, a pose-aware semantic exploration system that exploits a legged robot's intrinsic body pitch and roll with omnidirectional camera-LiDAR perception. The proposed pose-aware viewpoint sampling module selects body postures from partial object maps according to expected coverage gain, while aim-aligned execution reduces unnecessary body reorientation. Further, we introduce an object-centric viewpoint pruning strategy assisted by a vision-language model (VLM), which uses persistent observation history and bird's-eye-view (BEV) maps to reduce redundant inspection visits. The resulting semantic viewpoints are combined with geometric exploration viewpoints in a global exploration planner. Simulations show that POSE improves final target-surface coverage by 8-10 percentage points over the planar planning baseline while reducing exploration time by 17-32%, and achieves the highest mean object coverage AUC among the evaluated baselines. Real-world experiments with a legged robot carrying an omnidirectional camera-LiDAR suite in a machine shop further demonstrate the system's applicability. These results support adaptive body-posture planning for improving the coverage-efficiency trade-off in legged robot semantic exploration. We plan to release the code for community benefit in the future.