τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
2026-07-27 • Robotics
Robotics
AI summaryⓘ
The authors address the challenge of teaching robots to understand touch along with vision and language cues for better action control. They developed a framework called τ that learns how touch data relates to future visual scenes, helping the robot predict and act even with limited training data. They also created a new dataset named TacAura, combining touch, vision, and movement information from real manipulation tasks. Their approach improves robot performance and adaptability in handling new objects and situations.
tactile representationvision-language-action modelsspatiotemporal learningJoint-Embedding Predictive Architecture (JEPA)latent space supervisionproprioceptionsensor fusionrobot manipulationtouch sensing dataset
Authors
Ning Cheng, Jinan Xu, Wanlin Li, Yangzhi Chen, Jing Gao, Yiqun Wang, Kelan Peng, Wenjuan Han
Abstract
Learning the informative tactile representation while effectively adapting it to pretrained Vision-Language-Action (VLA) models remains challenging at both the data and modeling levels. At the data level, limited task-specific demonstrations constrain representation quality, whereas large-scale pretraining incurs substantial costs. At the modeling level, existing methods either focus on instantaneous contact states or model temporal interaction dynamics using 6D wrench sequences, leaving high-dimensional tactile signals underexplored. To address these challenges, we present τ, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation under limited data. This supervision operates in latent space and is used only during training, adding no deployment overhead. We also introduce TacAura, a dataset of synchronized vision, proprioception, and vision-based tactile signals across four representative contact-rich manipulation tasks. Experiments show that τ outperforms existing models and generalizes to unseen objects and scenes, delivering improved manipulation performance and robustness