Learning better object pose from RGB-D by linking shape and category cues
LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation
Computer Vision and Pattern Recognition
Summary
Figuring out exactly how an object is positioned in 3D from just one RGB-D image is tricky because you see only part of its shape. The authors developed LEGAU, a system that learns a kind of shape guide for each object category, which helps it guess the object's size, pose, and how visible parts match a standard shape. It combines color, depth, and category clues in one step using a transformer model. This joint learning leads to better results, especially when working with multiple categories or when switching from synthetic to real images.
What this means in practice
- •For robotics engineers: Improve robot handling by more accurately estimating position and orientation of diverse objects from a single RGB-D view using category-aware shape priors.
- •For augmented reality developers: Enhance object interaction by providing reliable 6D pose and size estimates for multiple object categories in real-time from RGB-D input.
Authors
Hongli Xu, Zhaowei Lu, Junwen Huang, Jiaqi Hu, Peter KT Yu, Benjamin Busam, Federico Tombari, Slobodan ilic
Abstract
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object pose and size, and a canonical Semantic Gaussian Field. Rather than treating reconstruction as a detached auxiliary task, LEGAU uses the Gaussian field as a category-conditioned structural prior that participates in multimodal feature fusion and provides global guidance for local pose reasoning. Conditioned on a categorical text embedding, LEGAU processes RGB-D observations through a transformer-based fusion module that integrates visual, geometric, and category-level cues, decoding the NOCS map, pose and size information and the Gaussian-based object representation. Extensive experiments on synthetic and real-world benchmarks show that this coupled pose-shape formulation achieves strong performance in a single-model multi-category setting, with up to 22\% on SOPE and competitive transfer to real-world data. These results highlight the benefit of jointly learning canonical correspondence, object shape, and pose alignment within a unified representation.