Visual geometry model improves robot precision in delicate assembly tasks

VGM-VS: Rethinking Visual Geometry Model for High-Precision Visual Servoing

RoboticsComputer Vision and Pattern Recognition

Summary

Robots need to know exactly where their tools and targets are to do precise jobs like plugging in cables or installing computer parts. The authors developed a method that lets a robot use pictures to figure out its position more accurately, even if parts of the target are hidden or hard to see. The robot learns from its own movements to understand the real scale of what it sees, so it can adjust precisely. Their method works fast and can handle tricky situations, helping robots do delicate assembly tasks with very high accuracy.

What this means in practice

  • For robotic assembly engineers: Enable robots to achieve submillimeter accuracy in cable and RAM installation tasks by improving pose estimation with visual geometry models and autonomous calibration.
  • For industrial automation teams: Improve real-time robot positioning during assembly with reliable visual feedback even under partial occlusions or poor texture conditions.

Authors

Yimin Pan, Sen Wang, You Zhou, Jianfeng Gao, Pengbo Sun, Ahmed M. Naguib, Zoltan-Csaba Marton

Abstract

We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image captured at the target configuration, we estimate the relative camera pose with a visual geometry model and apply it iteratively as the pose increment of a closed-loop pose-based visual servoing (PBVS) scheme. The geometry-aware representation acquired from large-scale pretraining keeps this estimate reliable when the target is occluded, weakly textured, or covers only a small part of the image. However, the scale ambiguity inherent to these models leaves the predicted translation defined up to an unknown scale, while the pose increment must be metric for robot control. We close this gap with a scene-specific metric adaptation: the robot autonomously records image--pose pairs along a predefined motion starting from the target pose, and we fine-tune the camera head on these data, jointly learning the hand--eye transform and thus removing the need for a dedicated calibration process. We evaluate our method on three real-world assembly tasks with demanding tolerances: USB-C cable picking, cable insertion, and RAM insertion. Running in real time at 30Hz, VGM-VS converges to submillimeter terminal accuracy on the cable tasks, and reaches success rates of 90--100\% when the target is moved during servoing. It converges in all trials under initial displacements of up to 30cm from the reference pose and with 50\% of the target object occluded, outperforming the compared visual servoing baselines.