SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation
2026-08-17 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors address a problem with Video Diffusion Transformers (DiTs), which use a slow and memory-heavy process called Self-Attention that gets much harder as videos get longer or higher resolution. They introduce SQuad, a new method that reduces the complexity of this process to a middle ground, making it much faster while keeping good quality. Instead of building a new model from scratch, they teach SQuad to imitate an already trained model through a two-step learning process. Their approach achieves similar video quality but runs much faster and uses less computing power.
Video Diffusion TransformersSelf-AttentionComputational complexityTokenSoftmaxDistillationFlow-Matching Supervised Fine-TuningDistribution Matching DistillationNeural Functional EvaluationsText-to-video generation
Authors
Animesh Karnewar, Denis Korzhenkov, Amirhossein Habibian, Mohsen Ghafoorian
Abstract
Video Diffusion Transformers (DiTs) spend most of their compute inside the Self-Attention operation, whose cost grows quadratically, $\mathcal{O}(n^2)$, with the number of latent tokens $n$. For the task of video generation, the token count is large, so this term dominates runtime and memory, and thereby caps the resolution and duration we can generate. Linear $\mathcal{O}(n)$ and low-rank $\mathcal{O}(nk)$ surrogates of Self-Attention trade the full softmax $QK^T$ for cheaper kernels, but rarely recover the original's expressivity, leaving a stubborn quality gap. Motivated by this, we propose SQuad, a Sub-Quadratic Attention Distillation framework that achieves a complexity of $\mathcal{O}(n\sqrt{n})$ in the resulting distilled Attention, naturally balancing the efficiency v/s expressivity trade-off. Instead of training our own Video DiT from scratch, which is prohibitively expensive, we fit a pretrained full softmax Self-Attention DiT into our proposed SQuad-Attention one by distilling the former in two stages: Flow-Matching Supervised Fine-Tuning (SFT), followed by improved Distribution Matching Distillation (DMD2) which additionally makes the sampling more efficient. On the Wan~2.2 5B text-to-video model, SQuAD matches the quadratic teacher on VBench ($83.20$ v/s $83.08$) while cutting the per-step per-block attention FLOPs by $\sim$$67\times$ and attention latency by $\sim$$11\times$, and end-to-end DiT latency by 2$\times$, all while also generating a video in only $6$ Neural Functional Evaluations (NFEs) instead of the default $100$.