Tactile signals improve video prediction of hand movements in manipulation

Dexterous Tactile World Model

Computer Vision and Pattern RecognitionMachine LearningRobotics

Summary

Predicting how hands will move to manipulate objects is hard just from video because touches and contacts are often invisible. The authors created a model that uses both video and touch information from gloves to better predict future hand motions. Their method predicts hand movements more accurately and especially improves predictions farther into the future. Adding touch information helps even when touch data is not available during prediction.

What this means in practice

  • For robotics engineers: Improve robotic manipulation by integrating tactile sensor data with video to more accurately predict and control hand-object interactions.
  • For augmented reality developers: Enhance hand interaction simulations by incorporating tactile feedback for more realistic and predictive modeling of hand movements.

Authors

Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan, Zhengxiang Yu, Fengyu Yang, Tianyu Liu, Zhiwen Fan, Daniel Rakita

Abstract

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.