Robot learns dexterous tasks using vision tactile and language data

STAR: Sparse Tactile Representation Learning in Vision-Tactile-Language-Action Models for Dexterous Manipulation

Robotics

Summary

Robots need to control many fingers and feel touch to handle objects well, but training them is hard because touch data is sparse and hard to use. The authors created a robot system to record a large dataset of videos, touch signals, and language commands during complex tasks with two hands. They developed STAR, a training approach that combines vision, touch, language, and action information to better understand sparse touch data. STAR helps robots successfully complete delicate tasks after some task-specific practice.

What this means in practice

  • For robotics engineers: Create robot control systems that use combined vision and touch data to improve complex hand manipulation tasks on robots.
  • For industrial automation teams: Train assembly line robots to perform precise multi-finger tasks more reliably using fused visual, tactile, and language data.

Authors

Xiangcheng Liu, Tianhao Wu, Le Zheng, Yidong Wang, Bowen Jiang, Mingjie Pan, Xinlin Ren, Yi Liu, Jianlan Luo

Abstract

Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.