Papers for

interactive media designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SAGE improves guidance stability in large generative models

SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models

Abstract: Diffusion Transformers with Mixture-of-Experts (MoE) routing are a leading recipe for scaling generative models. Classifier-Free Guidance (CFG) is essential for generation quality, yet excessively high guidance scales trigger collapse. We identify a previously unreported failure mode in their combination: the two CFG branches route independently, so their realized activations occupy different subspaces. The unconditional write then leaves the conditional subspace, and CFG amplifies that residual linearly in the guidance scale. We propose SAGE, a training-time regularizer that aligns unconditional MoE activations to the conditional subspace without restricting routing diversity, at zero inference cost. Toy experiments show that SAGE dramatically suppresses extreme drift by 9.2x. When scaled to a 1B-parameter text-to-image model, SAGE significantly improves generation quality, delivering a 9.3% boost in peak DPG-Bench performance. Extensive experiments demonstrate that SAGE consistently outperforms the baseline.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Generating images or other content with advanced AI models can sometimes fail when the system tries too hard to follow instructions, causing it to break down. The authors found that this happens because two parts of the AI model work in different ways and create conflicting signals. They created a new method called SAGE that trains these parts to work together better without slowing down the model. This results in steadier and higher-quality outputs when making images from text.
Open → 2609.34525v1

Text to image models speed up by adapting steps to prompt complexity

Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

Abstract: Text-to-image diffusion models often use a fixed number of denoising steps, balancing time costs and image quality. However, the optimal number of steps depends on the complexity of the input text prompt. We propose an adaptive diffusion controller that dynamically adjusts the number of steps to generate high-quality images efficiently, without additional model training. By leveraging a mixture of step schedules with varying step sizes and evaluating the error term discrepancy at each timestep, our method transitions between schedules to optimize performance. Experiments on COCO and DiffusionDB show that our approach reduces inference time while maintaining visual fidelity, offering a more efficient alternative for text-to-image diffusion models.

Tue 15 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Generating images from text usually takes the same fixed amount of time, even if some descriptions are simpler and don't need as much work. The authors created a method that changes how long the computer spends on each image, based on how tricky the description is. This makes image generation faster without losing picture quality. Their approach uses different step patterns and checks for errors during the process to decide when to switch. Tests show it works well on big image sets, saving time while keeping images looking good.
Open → 2609.16572v1