Constrained edit fields improve precision in text guided image editing

Constrained Edit Fields for Training-Free Flow Editing

Computer Vision and Pattern RecognitionMachine Learning

Summary

Editing images by following text instructions can accidentally change parts of the image that should stay the same. The authors designed a method called Constrained Edit Fields (CEF) to carefully identify which parts of the image need changing and which should not be touched. This approach uses information from the original image or from a rough edited version to decide where to focus edits. It helps keep background and unrelated areas more consistent while making the requested changes. Their method tested better than previous ones on a wide range of images and editing models.

What this means in practice

  • For digital artists: Use constrained edit fields to make precise image edits guided by text without unintentionally altering unrelated parts of the original image.
  • For mobile app developers: Integrate improved flow-based editing to provide users with more accurate and natural text-based photo modification features.

Authors

Jingxuan Kang, Yinsong Wang, Che Liu, Chen Qin

Abstract

Text-guided image editing aims to perform a desired edit while preserving source content unrelated to it. Pretrained rectified-flow models enable training-free editing of real images through modifications to their sampling trajectories. However, responses at locations unrelated to the desired edit can still accumulate along the editing trajectory and become visible in the final result. To overcome this, we propose Constrained Edit Fields (CEF), which assigns each spatial location a continuous edit responsibility that quantifies its relevance to the desired edit. CEF estimates edit responsibility directly from the source image when the relevant content is present. For edits whose target content is absent from the source, CEF first generates an unconstrained proposal to reveal its realized spatial support and then estimates responsibility from that proposal. At each editing step, CEF decomposes the base edit field into prompt-induced and trajectory-induced components, enabling edit responsibility to preserve instruction-relevant updates while suppressing unintended trajectory-induced changes. Evaluated on all 700 PIE-Bench examples, CEF achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment. On Stable Diffusion 3.5 Medium, it reduces these metrics over the previous best results by 10.2%, 21.2%, and 48.0%, respectively.