SAGE improves guidance stability in large generative models
SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models
Computer Vision and Pattern Recognition
Summary
Generating images or other content with advanced AI models can sometimes fail when the system tries too hard to follow instructions, causing it to break down. The authors found that this happens because two parts of the AI model work in different ways and create conflicting signals. They created a new method called SAGE that trains these parts to work together better without slowing down the model. This results in steadier and higher-quality outputs when making images from text.
What this means in practice
- •For ai model developers: Enhance large-scale image generation by integrating SAGE for more stable and higher-quality outputs with no added runtime cost.
- •For interactive media designers: Create more reliable text-guided visuals in interactive applications by reducing failures caused by guidance scale issues using the SAGE technique.
Authors
Boyu Zhang, Yangming Cheng, Ning Zhang, Pengfei Liu, Weijie Li, Yifan Gao, Hangyu Li, Litong Gong
Abstract
Diffusion Transformers with Mixture-of-Experts (MoE) routing are a leading recipe for scaling generative models. Classifier-Free Guidance (CFG) is essential for generation quality, yet excessively high guidance scales trigger collapse. We identify a previously unreported failure mode in their combination: the two CFG branches route independently, so their realized activations occupy different subspaces. The unconditional write then leaves the conditional subspace, and CFG amplifies that residual linearly in the guidance scale. We propose SAGE, a training-time regularizer that aligns unconditional MoE activations to the conditional subspace without restricting routing diversity, at zero inference cost. Toy experiments show that SAGE dramatically suppresses extreme drift by 9.2x. When scaled to a 1B-parameter text-to-image model, SAGE significantly improves generation quality, delivering a 9.3% boost in peak DPG-Bench performance. Extensive experiments demonstrate that SAGE consistently outperforms the baseline.