Diffusion transformers improve image quality with timestep aware gating

TSGate: Timestep-Aware Gated Attention for Diffusion Transformers

Computer Vision and Pattern Recognition

Summary

Diffusion Transformers are tools used to create images from text prompts, but they struggle when given unusual or unexpected descriptions. The authors found that problems happen because the model’s attention to the prompt weakens early in the image-making process. They introduced a new method called TSGate that adjusts how attention gates work during different stages, leading to better image generation even with uncommon prompts. Experiments show this method improves performance compared to previous approaches.

What this means in practice

  • For ai image generation teams: Enhance image quality in text-to-image models by improving prompt adherence and handling free-form user inputs better during generation.
  • For video game developers: Improve in-game content creation tools that generate high-quality images or scenes from player text input by reducing quality drops with unusual prompts.

Authors

Boyu Zhang, Yifan Liu, Shuxia Lin, Qingjian Ni, Yinfei Xu, Xu Yang

Abstract

Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-of-domain (OOD) prompts, including the free-form descriptions supplied by users at inference time. Although LLM-based rewriting can convert these prompts into structured formats, it does not guarantee that the rewritten prompts align with the training distribution. Our analysis links this degradation to attention sinks and reduced early-step image-to-text attention and shows that sink suppression alone is insufficient to restore generation quality. Despite effective sink suppression, models trained with standard gated attention exhibit reduced early-step image-to-text attention and suboptimal generation quality. Based on these insights, we propose Timestep-Aware Gated Attention (TSGate), which injects a timestep-conditioned bias into the gate signal so that gating behavior adapts across denoising steps. Extensive experiments show that TSGate consistently outperforms both the baseline and standard gated attention across multiple benchmarks, improving the raw-prompt DPG score by 9.5% over the baseline.