Summary
Video diffusion transformers, which create videos by gradually refining images, spend a lot of time calculating attention, a process that figures out how parts of the video relate to each other. The authors address this by improving block-sparse attention, which skips unnecessary calculations by focusing only on important blocks defined by a logical mask. They introduce Tessera, a system that separates how these blocks are logically defined from how they actually run on GPUs, making it adaptable to different GPU types and attention patterns. Tessera organizes work on the GPU efficiently and picks execution plans quickly, resulting in big speed improvements for video generation models.
What this means in practice
- •For video ai engineers: Use Tessera to accelerate attention computations in video diffusion models on various NVIDIA GPUs, improving inference speed without changing model masks.
- •For graphics hardware developers: Adapt Tessera’s approach to optimize GPU execution strategies that decouple logical task layouts from hardware specifics, enhancing flexible performance tuning.
Authors
Shanghao Liu, Xiaoyun Yu, Wanting Li, Wenqi Jiang
Abstract
Attention computation makes inference expensive in video diffusion transformers (vDiTs), which generate videos through iterative denoising. Block-sparse attention (BSA) reduces this cost by computing only blocks selected by a logical mask, which specifies attention interactions to compute. However, coupling logical block geometry to execution choices limits adaptation to varying masks and graphics processing units (GPUs), while runtime kernel specialization can incur preparation overhead that outweighs execution time savings. We present Tessera, a specialized runtime for dynamic BSA that decouples logical masks from GPU execution while preserving specified attention interactions. Its physical mapping layer retains, combines, or subdivides logical attention blocks into physical tiles suited to different attention mask shapes and GPU architectures. Its task organization layer groups and schedules tiles within GPU tasks to reuse data, expose parallelism, and overlap data movement with computation. Finally, profile-guided regime selection enables low- overhead execution plan selection through a lookup table constructed from offline profiling. We implement Tessera with specialized CUDA kernels supporting four NVIDIA GPU generations. Evaluated on 2,315 real attention masks and industrial video diffusion models, Tessera achieves up to 6.79x BSA request speedup over baseline systems in the evaluated video diffusion models.