World action models improve generalization by preparing future data

What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

Computer Vision and Pattern RecognitionRobotics

Summary

Predicting future events helps computer models learn better, but doing this during actual use can be very costly. The authors found that ignoring future data during use causes the model to lose important generalization abilities, even if it works fine on familiar tasks. They discovered that simply preparing the future data once, rather than fully generating it each time, keeps the benefits while saving on computation. Their method, Simple-WAM, balances accuracy and efficiency better than previous approaches.

What this means in practice

  • For robotics engineers: Use simplified future data preparation to improve robot action prediction under new or changing environments.
  • For autonomous vehicle developers: Integrate efficient future modeling methods to enhance decision making in varied driving scenarios with limited data.

Authors

Renping Zhou, Zanlin Ni, Zihao Fan, Guohao Fu, Zeyu Liu, Hao Shi, Jie Zhang, Chi Bene Chen, Yang Yue, Xueyang Fu, Gao Huang

Abstract

World action models (WAMs) predict the future alongside actions during \emph{training}. Due to the heavy computation cost of video denoising, whether the future must still be generated during \emph{inference} is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration. We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: \emph{environmental perturbation}, \emph{data efficiency}, and \emph{task generalization}. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from \emph{preparing} the future, not \emph{generating} it. We therefore propose \textbf{Simple-WAM}, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs. Project Page: \href{https://zrporz.github.io/Simple-WAM-Web/}{\textcolor{panton}{\texttt{https://zrporz.github.io/Simple-WAM-Web}}}