Adaptive joint attention speeds up conditional image generation with diffusion transformers
RefAdapt-DiT: Adaptive Joint Attention for Reference-Conditioned Diffusion Transformers
Computer Vision and Pattern Recognition
Summary
Generating images with special AI models called Diffusion Transformers is usually slow when these models need extra information to create images, because they process reference details many times. The authors found that parts of this reference information often change slowly and are sometimes ignored by the model, meaning some computations are unnecessary. They designed a new method called RefAdapt that smartly decides when to update this reference information based on how much the model is paying attention to it, without retraining the AI. This approach speeds up the image generation process significantly while keeping similar image quality.
What this means in practice
- •For image generation developers: Reduce computation time in conditional image generation systems by adaptively controlling attention to reference inputs for faster outputs.
- •For image editing software teams: Accelerate few-step image editing tasks that rely on diffusion transformers by adaptively managing reference updates without losing output quality.
Authors
Jian Tang, Jiawei Fan, Qiannan Zhou, Qingbin Liu, Jiang Bian, Zang Li
Abstract
Diffusion Transformers (DiTs) have become the standard backbone for high-quality generative modeling, yet deploying them in conditional generation tasks remains computationally prohibitive because bidirectional joint attention repeatedly processes large reference streams. While existing optimization schemes mitigate generic temporal redundancy, they typically rely on coarse-grained static reuse and overlook the distinct dynamics of references and targets. Specifically, we observe that reference representations often evolve slowly along the generation trajectory, while the target often assigns little attention mass to them; reference drift and this target-to-reference exposure jointly shape how strongly stale reference states affect the target. To exploit these patterns, we introduce \RefAdapt, a training-free framework for adaptive control of joint attention between references and targets. Instead of rigid static strategies, \RefAdapt combines consecutive target-Q change with previously observed target-to-reference attention mass to control reference computation adaptively at block granularity. Under ultra-few-step settings, \RefAdapt enables speedups of up to $2.097\times$ on 4-step MiniMax H3 and $3.54\times$ on 8-step Qwen Image Edit, while maintaining comparable visual quality.