Vision combined with touch improves cluttered dexterous robot grasping

When Does Touch Matter? Charting the Vision-Interaction Gap in Cluttered Dexterous Grasping

Robotics

Summary

Robots need to grasp objects skillfully even when things are messy and overlapping. The authors found that using touch sensors together with vision helps robots grab objects more reliably, especially when vision alone fails due to clutter or hidden parts. Their study shows that touch feedback allows robots to detect bad grips earlier and adjust before lifting objects. This work is the first to test and compare how vision and different types of touch sensing work together in real-world cluttered grabbing tasks.

What this means in practice

  • For industrial robot integrators: Improve object picking in cluttered factory spaces by combining visual and touch sensors for more reliable robot grasping.
  • For warehouse automation teams: Design robot grippers that use touch sensing to handle tightly packed items safely and efficiently, reducing errors in sorting tasks.

Authors

Hao Jiang, Luis Dominguez, Daniel Seita

Abstract

Dexterous grasping in clutter poses a basic sensing question: when do tactile measurements and external wrench estimates improve on visual geometry? Occlusion and contact can obscure grasp quality, motivating a controlled evaluation of these interaction signals. We present a controlled real-world study over five tabletop scene conditions on a dexterous system that combines vision, per-finger and wrist wrench estimates, and distributed fingertip taxels. With demonstrations, visual observations, action space, and compliant control fixed, we compare vision-only, wrench, taxel, and combined policies plus representation and fusion baselines. The combined policy succeeds in 24/25 trials versus 14/25 for vision only, and 15/15 versus 6/15 across the three confined conditions. Ablations show that wrench and taxel feedback are complementary. Behavioral comparisons show that interaction feedback enables earlier rejection of inadequate contacts, regrasping before lift, and more stable grasps. To our knowledge, this is the first real-world study to combine and separately evaluate these interaction modalities for target-oriented dexterous grasping in clutter. These results chart a widening vision-interaction gap and position cluttered dexterous grasping as a benchmark for determining when the learned policy needs interaction sensing. Project website: https://interaction-dex-grasp.github.io/