World synesthesia model improves robot hand manipulation of many objects
WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation
Robotics
Summary
Manipulating objects inside a robot’s hand is hard because the robot needs to understand the object’s shape, position, and feeling from limited and noisy data. The authors propose WM-Craftnet, a system that learns from multiple sensor types to create a clearer understanding of the object’s state. This helps the robot handle various objects more reliably, even when things change or get disturbed. Their model also cleans up noisy depth sensor data to improve real-world performance. Training on some objects allows it to quickly adapt to many others, improving the robot’s dexterity and robustness.
dexterous in-hand manipulationworld modelproprioceptiondepth sensingtactile sensingreinforcement learninglatent dynamicssim-to-real transfermulti-object manipulation
Authors
Jie Yin, Zeyuan Zhao, Xiaojing Tan, Yang Liu, Chiyu Wang, Xinyang Gu
Abstract
Generalizable and robust dexterous in-hand manipulation requires a policy to infer object pose, geometry, contact, and potential slip from partial and noisy observations. Although recent tactile and visuotactile RL methods achieve strong in-hand rotation in controlled settings, their robustness often degrades under pose shifts, force disturbances, and object variation. We propose WM-Craftnet, a world-model-conditioned framework that learns compact action-conditioned latent dynamics from proprioception, depth, tactile sensing, and actions, supervised by multimodal reconstruction and reward prediction. Rather than using the world model for latent imagination or policy optimization, WM-Craftnet uses the learned World Synesthesia Model (WSM) as recurrent task context for an asymmetric actor--critic policy. Importantly, WSM is trained to reconstruct clean depth targets from noisy depth inputs, providing a denoised geometric state for real-robot deployment. Ablations over recurrent baselines, auxiliary heads, tactile masking, and WSM modality heads show that predictive world modeling, clean-depth supervision, and tactile contact cues all shape the learned state. A WSM pretrained on nine \(z\)-axis objects serves as a reusable prior for \(49\)-object downstream policy learning. This context improves multi-object rotation, with quantitative and qualitative evidence for unseen-object, perturbation-recovery, and sim-to-real transfer.