DecFlowEdit improves image edits by better separating changes from background

DecFlowEdit: Self-Localized Flow-based Image Editing via Guidance Decoupling

Computer Vision and Pattern Recognition

Summary

Editing images without messing up the background can be hard because changes sometimes spill over where they shouldn’t. The authors found that the usual way these edits guide the changes causes unwanted background effects. They developed DecFlowEdit, a technique that separates how it figures out where to edit from how it applies the edits. This approach keeps the background intact while still making good changes, and it works without extra training or complicated tools. Tests showed it preserves backgrounds much better than previous methods while keeping edits accurate.

What this means in practice

  • For graphic designers: Enhance image editing tools to make precise object changes without affecting surrounding background areas.
  • For mobile app developers: Integrate flow-based editing capabilities that improve background preservation in photo editing apps without extra user input.

Authors

Zheyuan Zhan, Can Wang, Jiawei Chen, Chun Chen, Siwei Lyu, Zeyu Zheng, Defang Chen

Abstract

Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial background leakage. Matching these guidance scales, for example by removing CFG, improves edit-relevant localization but severely degrades editability. To get the best of both worlds, we propose DecFlowEdit, which decouples the optimal guidance scales for localization and for editing in flow-based generative models. In particular, DecFlowEdit first extracts an edit-relevant prior by temporally aggregating velocity differences evaluated without CFG, and then uses this prior to reweight the original updates under default CFG. Our method remains training-free and inversion-free, requiring neither external spatial masks nor attention manipulation. Experiments on PIE-Bench across FLUX, SD3, and SD3.5 show that DecFlowEdit improves background preservation, reducing structure distance by approximately 61 to 73 percent and background LPIPS by 68 to 80 percent relative to FlowEdit at comparable editing fidelity.