Correcting attention improves fidelity in diffusion visual editors
Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing
Computer Vision and Pattern Recognition
Summary
Diffusion-based visual editors sometimes have trouble closely following the details of user-provided reference images when making edits. The authors found that this is because the editing process often pays very little attention to the reference images. They created RefGAP, a way to adjust how much attention is given to references during editing without extra training. This adjustment helps the editor better copy important details from the reference, improving results like face or head swapping in images and videos.
What this means in practice
- •For image and video editing developers: Enhance identity preservation when performing edits like face and head swaps in diffusion-based visual editing tools.
- •For fashion technology teams: Improve virtual try-on systems by better aligning clothing references to user images during editing.
Authors
Yanan Wang, Shengcai Liao, Guangyi Liu, Xiaodan Liang
Abstract
Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. We introduce RefGAP, a training-free correction that determines logit-offset magnitudes online at each layer from the reference-attention mass measured during the forward pass. Positive offsets to reference logits strengthen reference usage by edit-region queries, while negative offsets for keep-region queries limit reference-induced changes outside the edit. Two global coefficients control the correction; they are selected once on validation data from four development diffusion editors and held fixed. Across seven diffusion-based image/video editors, RefGAP improves identity fidelity in head swapping and face swapping. RefGAP achieves a fidelity-preservation trade-off comparable to separately tuned constant edit-side biases, without per-approach strength sweeps. Additional experiments on virtual try-on and background replacement evaluate transfer beyond identity editing.