Efficient attention helps simulate interactive video worlds longer
WorldAttention: An Efficient Attention Architecture for Interactive Video World Models
Computer Vision and Pattern Recognition
Summary
Simulating interactive video environments guided by text is challenging when trying to remember everything that happened before without slowing down. The authors created WorldAttention, a new way of organizing memory and attention that mixes different techniques to keep track of long histories efficiently. This lets systems remember more about past video frames without using too much computer memory or processing power. Their tests show that WorldAttention better maintains consistency in long video simulations than earlier methods.
What this means in practice
- •For game developers: Build interactive video game worlds that respond to player text instructions over long sessions without memory slowdowns.
- •For robotics engineers: Enable robots to simulate and plan actions in complex video environments using detailed long-term memory of past events.
Authors
Zeyu Zhang, Jinyuan Mao, Dakai An, Wangbo Zhao, Hanfeng Lu, Jiasheng Tang, Yinghao Yu, Wei Wang, Bohan Zhuang
Abstract
Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current frameworks primarily rely on sliding-window mechanisms to bound computational complexity. However, this approach inherently sacrifices historical context, undermining the long-range interactive capabilities. Conversely, maintaining a full-history cache remains computationally prohibitive and memory-intensive: the quadratic complexity of attention leads to excessive computational overhead, while the linear growth of the KV cache inevitably leads to GPU memory saturation. To overcome these limitations, we propose WorldAttention, a system-oriented attention architecture that achieves high efficiency through the co-design of specialized attention kernels and hierarchical KV cache management. First, we introduce Hybrid Sparse Attention (HSA), which integrates linear global attention supplemented with head-adaptive sparse attention. Additionally, we design a Hierarchical KV Cache (HKV) that organizes historical KV pairs into semantically indexed pages across multi-tier memory, enabling fine-grained retrieval and controlled GPU residency. These two designs are supported by tailored kernels to effectively translate their theoretical efficiency into real-world performance. Extensive experiments on VBench-Long and InterVBench demonstrate that WorldAttention consistently surpasses prior state-of-the-art methods, achieving subject consistency scores of 0.9472 on VBench-Long and 0.9668 on InterVBench, respectively.