Video sparse attention speeds up generation with fine-grained routing

Improving Video Sparse Attention with Fine-grained Router and Sparse Rebasing

Computer Vision and Pattern Recognition

Summary

Processing videos efficiently is challenging because it requires lots of computations to focus on important parts. The authors developed VSA2, a method that smartly picks which parts of the video to pay attention to, using a ‘fine-grained router’ that adjusts how much is processed for each part. They also found a special training approach that helps models perform even better when they focus less, making video generation faster without losing quality. Their method can be used during different training stages and speeds up video creation significantly compared to previous techniques.

What this means in practice

  • For video processing engineers: Implement faster video generation pipelines by replacing full attention with VSA2’s sparse attention to reduce compute with maintained or improved quality.
  • For machine learning infrastructure teams: Optimize training workflows for video transformer models by integrating VSA2’s dynamic sparse attention and curriculum training to reduce resource usage.

Authors

Peiyuan Zhang, Guoqiang Wei, Yilong Zhao, Zixiang Zhang, Wei Zhou, Will Lin, Heng Zhang, Xiaonan Nie, Yan Zeng, Hao Zhang

Abstract

We present VSA2, a frontier trainable sparse attention for video DiTs. VSA2 includes a variety of new architectural features and training procedures that we apply across all stages of the DiT development cycle, including pretraining, RL, and inference, to produce a DiT with comparable or better quality than a full attention counterpart. Architecturally, VSA2 introduces a fine-grained router that improves the precision of identifying critical tokens and supports dynamic computation by allowing each query to attend to a variable number of key-value pairs. In training, we identify a Hard-to-Easy Curriculum, where models trained under high sparsity and later evaluated with lower sparsity during inference not only generalize effectively, but also outperform models trained with full attention in motion quality. VSA2 is also flexible: it can replace full attention during the middle of progressive low-to-high resolution pretraining, rebasing early-stage full-attention checkpoints. Experiments show that VSA2 reduces attention computation by half over VSA with lower loss. On 720p videos, it accelerates attention by 8.9x and end-to-end generation by 4.62x compared to the FlashAttention-3 baseline, while achieving comparable or better video quality.