Egocentric stereo model improves 3D hand tracking in real world

ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Seeing and understanding hand movements in 3D from wearable cameras is hard, especially outside controlled labs. The authors created ESTHER, a model that uses two cameras worn on the head to better estimate hand positions and shapes in real time. They also built a new dataset to help train and test such models in natural environments. Their approach is more accurate, works well even when one camera view is lost, and keeps real-world size scales right, unlike prior methods.

What this means in practice

  • For augmented reality developers: Enable accurate and robust 3D hand tracking for AR glasses in real-world environments using readily wearable stereo cameras.$Commercial implications: Makes precise hand interaction features possible on consumer AR devices by reliably capturing 3D hand motions outdoors.
  • For robotics engineers: Provide robotic systems with human-like hand perception from egocentric stereo vision to improve manipulation tasks in natural settings.

Authors

Hongyu Ma, Hairong Qu, Shiqi Zhao, Yongsong Yang, Peng Yin

Abstract

Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-end model nor an in-the-wild benchmark. We propose ESTHER, a model whose stereo geometry, temporal reasoning, and output representation are designed for wearable egocentric stereo. It is trained on pseudo-labels from a calibrated labeling pipeline and in turn assembles our benchmark ESTHER3D, an egocentric stereo hand dataset pairing a large in-the-wild training set of model-generated labels with a motion capture test set of true metric ground truth. Experiments show state-of-the-art accu?racy, superior external generalization, and robustness to the missing views, dropped frames, and lighting and motion blur extremes of real egocentric capture that break existing meth?ods. This robustness runs deeper than graceful degradation: stereo guidance teaches the model to bind apparent hand scale to metric depth, so it not only adapts to different stereo rigs and modalities with minimal fine-tuning, but more strikingly preserves true metric scale even after collapsing to a single monocular view.