ViHaTeleop: A Low-Cost, Lightweight Visual-Haptic Teleoperation System for Dexterous Manipulation Learning
2026-08-17 • Robotics
Robotics
AI summaryⓘ
The authors developed ViHaTeleop, an affordable and lightweight system that helps people control robot hands using visual and touch feedback. They used special cameras and vibrating motors to improve how well people can perform delicate tasks with the robot. In tests with nine people doing six tricky tasks, adding touch feedback helped them succeed more often. They also created a simulation tool that uses these visual and touch signals to train robots, showing better results when robots use touch cues. Overall, the authors show that combining sight and touch helps robots learn delicate manipulations from humans.
Learning from demonstrationTeleoperationSLAMHaptic feedbackVibrotactile feedbackDexterous manipulationSimulationVisual-tactile policyLEAP HandPeg-in-hole task
Authors
Fucai Zhu, Yanhou Lai, Paul Maestre, Koichi Hashimoto
Abstract
Learning from demonstration is a promising approach for dexterous manipulation, but collecting high-quality contact-critical demonstrations remains difficult with low-cost teleoperation hardware. We present ViHaTeleop, a lightweight (0.7 kg), low-cost (\$550) visual-haptic teleoperation system with SLAM-based wrist tracking, camera-based hand tracking, and finger-wise vibrotactile feedback through Linear Resonant Actuators (LRA). The system includes several design choices (LED illumination, fisheye hand camera, and tactile-aware retargeting constraints) and is deployed on Franka + LEAP Hand + 9DTact in both real and simulated environments. Under matched with/without-haptic conditions with nine participants across six contact-critical tasks, haptics improved success rates across all tasks (+2.2 to +15.6 percentage points), while completion-time effects were task-dependent. Subjective ratings showed significant gains in contact clarity and grasp confidence in both simulation and real-world settings (Wilcoxon signed-rank, $p<0.05$). We also integrate a lightweight depth-camera-based tactile proxy in Isaac Sim, enabling a full pipeline from multi-modal demonstration collection to visual-tactile policy training. Preliminary downstream validation by training visual-tactile policies from collected demonstrations shows tactile cues benefit contact-critical subtasks (peg-in-hole: +17 percentage points over vision-only).