Whole-body tactile sensing boosts humanoid robot control success
Uni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-Manipulation
Robotics
Summary
Robots that walk and manipulate objects often have trouble sensing physical contact just by looking or feeling their joints. The authors show that adding detailed touch sensing over the entire robot body helps it better understand and react to its surroundings. Their method predicts future touch and vision signals to create a more complete picture for robot control. Tests on real robots show this approach works much better than controlling robots without whole-body touch sensing.
What this means in practice
- •For robotics engineers: Improve contact-rich control of humanoid robots by integrating whole-body tactile sensing with vision and proprioception.
- •For industrial robot integrators: Enhance reliability of robots performing complex tasks requiring physical contact with humans or objects using predictive tactile learning.
Authors
Zihao Wang, Shutong Liu, Siqi Zheng, Liu Cao, Ruoqi Chen, Rundong Liu, Yanchao Yang, Mengdi Xu
Abstract
Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions, distributed tactile sensing preserves spatially resolved contact patterns across the robot body. We therefore study how to integrate such whole-body tactile information into vision-language-action (VLA) policies for contact-rich control. Our approach, Uni-VLaT, introduces a tactile pathway whose latent state is trained not only for action generation, but also to predict future tactile, proprioceptive, and visual representations. This predictive objective builds a tactile-anchored multimodal context, encouraging a more structured understanding of the physical world. We evaluate Uni-VLaT on five real-robot tasks covering tactile-triggered locomotion, sustained physical interaction, human-robot contact, and loco-manipulation. Uni-VLaT achieves a 75% average success rate, outperforming a baseline without tactile input by 43 points and a tactile-input baseline without predictive supervision by 7 points. Across two pretrained VLA backbones, our method improves Table Sweeping by 30 points on both backbones and Back-Tap Walking by 85-90 points. Ablations further show that contextualized tactile prediction and absolute future targets are critical to performance. These results indicate that predictive tactile learning provides an effective route for extending pretrained VLA policies to whole-body physical interaction.