World action models improve robot and human action prediction accuracy

From World Models to World Action Models: Rethinking Next-State Prediction

Robotics

Summary

Predicting what will happen next is important for robots and AI to understand and interact with their surroundings. The authors found that fixing how the future is predicted limits learning, so they made a method that looks at many different ways to describe the same future scene. This helps the AI learn better how actions lead to changes, even when used by both humans and robots with different appearances. Their approach improved performance in experiments, showing it helps AI control systems learn more effectively.

What this means in practice

  • For robotics engineers: Use dynamic next-state projections to improve robot control policies by better capturing diverse action outcomes.
  • For human-robot interaction teams: Translate human movement experiences into robotic policy gains by aligning diverse next-state representations.

Authors

Tingyu Yuan, Ziming Ji, Biaoliang Guan, Wen Ye, Wenrui Tian, Zhaopeng Gu, Feihong Zhang, Xu Yang, Yan Huang, Zhaowen Li, Chaoyang Zhao, Jinqiao Wang

Abstract

Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.