Winning method predicts hand object contact points from stereo images

GOLF: Global Observation with Local Focus for Calibration-Aware Stereo Interaction Field Estimation

Computer Vision and Pattern Recognition

Summary

This work focuses on figuring out exactly where a person's hand touches an object in 3D space, using paired images taken from slightly different viewpoints. The team combined big-picture scene information with detailed close-up data about the hand and object, using special geometric math that helps relate lines and points in space. They used an advanced type of AI model called a Vision Transformer that was carefully adapted for this task. Their method was tested in a competition and achieved the best results for predicting the nearest points on an object to specific hand joints.

stereo vision3D vector estimationhand-object interactionPlücker ray geometryVision TransformerLoRALayerNormegocentric cameramean ADEensemble learning

Authors

Minqiang Zou, Riqiang Jin, Zhi Lv, Dong Luo, Lianghai Tian, Zhenyu Zhao, Qi Xu, Tong Wu, Mochen Yu, Yao Tang

Abstract

We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2026. Given synchronized egocentric stereo views, the task is to predict a 3D vector from each of 21 hand joints to the closest point on the manipulated object. GOLF combines dense global context, locally sampled hand/object evidence, and common-frame Plücker-ray geometry. We adapt DINOv3 ViT-H+/16 with LoRA and trainable LayerNorm parameters, then jointly decode both interaction fields. Our primary model achieves an official score of 27.61 and a mean ADE of 27.96 mm on the hidden test set. An equal-weight ensemble with a complementary directly fine-tuned variant improves these results to an official score of 27.47 and a mean ADE of 27.82 mm, securing first place.