InfiniHand streams accurate hand motion tracking from first-person video
InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
Computer Vision and Pattern Recognition
Summary
Tracking how hands move in 3D from video recorded by a camera worn on the head is hard because the camera moves too. The authors present InfiniHand, a way to estimate a hand’s 3D shape and position along with the camera’s motion all at once, without needing separate tools for each. They trained InfiniHand with a very large set of videos and showed it works better and faster than previous methods. This makes it easier to follow hand movements accurately even while the camera wearer is moving around.
What this means in practice
- •For virtual reality developers: Integrate a faster and more accurate hand tracking system for immersive VR experiences using egocentric cameras.$Commercial implications: Enables commercial VR products with improved hand motion capture using standard head-mounted cameras.
- •For robotics engineers: Use streaming hand and camera pose estimates to improve human-robot interaction through real-time gesture understanding.
Authors
Kerui Ren, Kaiwen Song, Weiguang Zhao, Yuxi Wang, Yufei Liu, Bo Dai, Haoyu Guo, Chunhua Shen, Mulin Yu, Tao Lu, Junting Dong
Abstract
World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhead. To address these limitations, we present InfiniHand, an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video. InfiniHand integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture. We train InfiniHand in two progressive stages by first learning robust camera-space hand priors and then extending to streaming world-space reconstruction. To support this process, we aggregate a pretraining corpus of approximately 5,000 hours of egocentric data across multiple public datasets. Extensive evaluations demonstrate that InfiniHand outperforms state-of-the-art baselines on in-domain benchmarks, achieving a 21.4% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift. Furthermore, InfiniHand generalizes robustly to in-the-wild videos and operates at 11.19 FPS, delivering more than twice the throughput of HaWoR.