World-action models predict future movements using multiple visual data types
Modality-Autoregressive World-Action Models
Robotics
Summary
Predicting what will happen next is hard when only looking at regular color images. The authors show that using different types of visual information, like depth or motion points, helps computers better guess future observations and actions. They created a new model, ModAR, that predicts these different data types in sequence, improving performance without lots of extra training. This model does better at tasks involving two-handed actions and learns even more from watching videos made by people.
What this means in practice
- •For robotics engineers: Improve robot task planning by using multiple visual signals to predict future states and actions with fewer training requirements.
- •For industrial automation teams: Enhance controllers for complex two-handed assembly by integrating sequential predictions of motion and depth features alongside visual inputs.
Authors
Adam Hung, Bardienus P. Duisterhof, Deva Ramanan, Jeffrey Ichnowski
Abstract
World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalities within WAMs remains an open question. We introduce ModAR, the first WAM to autoregressively denoise multiple future modalities before predicting actions. This allows each prediction to condition on previously generated modalities. We train from scratch to systematically study how training-data mixtures, predicted modalities, and WAM formulations affect performance. In our evaluations, WAMs benefit from predicting point tracks, DINO features, and depth maps, while additionally predicting future RGB does not provide a consistent benefit. We also find that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales. We also fine-tune the video-model-initialized WAM Flex-$π$ on the same data; ModAR achieves a slightly higher observed average success rate (75% vs. 72%) while using approximately $20\times$ fewer training FLOPs and no pretraining. On three real-world bimanual tasks, ModAR outperforms baselines and improves with human videos.