WorldPlay2 improves real-time interactive world models with better control and memory

WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon

Computer Vision and Pattern Recognition

Summary

Interactive world models are computer programs that simulate virtual environments, needing to respond instantly to different controls while keeping the scene consistent over long times. The authors present WorldPlay2, which combines a new way to handle different types of controls with smart memory compression and stable training techniques. These improvements let the system respond quickly, understand complex commands about characters and scenes, and keep the environment steady for longer. Tests show WorldPlay2 works well across many scenarios and beats earlier methods in performance.

What this means in practice

  • For game developers: Use enhanced control and memory techniques to create more responsive and consistent game worlds where player actions affect complex scenes over time.$Commercial implications: Enables new interactive gaming experiences with real-time complex scene control and stable long-term environment consistency, attractive for game publishers and studios.
  • For virtual reality developers: Build virtual environments that react instantly and stay consistent during extended interactions, improving immersion and user experience.

Authors

Haiyu Zhang, Wenqiang Sun, Tengfei Wang, Junta Wu, Jun Zhang, Yunhong Wang, Yu Qiao, Chunchao Guo

Abstract

Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time responsiveness. In this paper, we present WorldPlay2, an interactive world model that couples a factorized hybrid control interface with a co-design of compressed memory and stable distillation. 1) Our factorized hybrid control interface integrates frame-aligned action control with structured semantic control that explicitly disentangles scene appearance, character identity, and dynamic semantic events, thereby facilitating effective control learning. 2) To achieve efficient long-horizon modeling, we compress historical contexts into compact memory tokens shared by the autoregressive student and the bidirectional teacher. This design enables clip-wise, memory-conditioned score evaluation instead of jointly processing an entire long rollout, substantially reducing distillation overhead. 3) We further propose Stable Forcing, which initializes the autoregressive student via a few-step strategy and leverages full-rollout replay to preserve the quality of long-horizon rollouts, ensuring robust and stable distillation. Extensive experiments demonstrate the strong generalizability of our model and its superior performance compared to existing methods.