World action models improve robot learning with calibrated representations
Rethinking Representations for World-Action Modeling
Computer Vision and Pattern RecognitionRobotics
Summary
Teaching robots to act and predict what they will see next is tricky because it depends on how the robot understands the world. The authors found that just making the robot’s visual predictions very accurate or using pre-made visual features doesn't guarantee good actions. They created a new method called ReWAM that organizes these visual features in a way that helps the robot learn better policies by letting the robot’s actions shape what it pays attention to. This approach improved robot task success on two robot training setups without needing extra video training.
What this means in practice
- •For robotics engineers: Train robot control policies that better integrate perception and action without costly video generative pre-training, improving real-world task success.
- •For autonomous system developers: Design compact state representations from visual features to improve prediction of future states and actions in embodied AI applications.
Authors
Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang
Abstract
World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.