Unified image editors improve quality by switching text generation mode

When Text-to-Image Helps Editing: The Effects of Conditioning During Denoising

Computer Vision and Pattern Recognition

Summary

Image editing models often keep using the original image as a guide while making changes, but the authors wondered if letting the models borrow abilities from text-to-image generation could help. They found that switching the model to focus on generating images from text for parts of the editing process improved how good the final edits were, without losing much of the original image’s features. This balance depends on when the model switches between using the original image or text instructions. So, mixing both ways the model was trained helps produce better edits.

What this means in practice

  • For digital artists: Create more accurate and faithful image edits by timing the switch between original image and text guidance during editing.
  • For photo editing software developers: Improve editing tools by integrating controlled switching between image and text-based model conditioning to enhance edit quality while preserving original details.$Commercial implications: Enables advanced editing features that produce higher-quality outputs, offering competitive advantages in photo editing products.

Authors

Lidia Troeshestova, Alexander Ustyuzhanin, Sergey Kastryulin

Abstract

Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-image conditioning throughout denoising. We ask whether editing can benefit from T2I, and study how the effects of conditioning vary across edits and denoising stages. In pure editing, source attention declines for some edits over the sampling trajectory. This observation led us to task switching, which lets the model draw on its T2I capabilities. Across three unified editors and four benchmarks, switching to the T2I task for bounded intervals improves edit quality, while mean perceptual preservation remains close to pure editing across all three models. Unified editors therefore benefit from using both conditioning modes they are trained for, and the timing of the switch sets the balance between quality and preservation.