Papers for

image and video editing developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Correcting attention improves fidelity in diffusion visual editors

Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing

Abstract: Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. We introduce RefGAP, a training-free correction that determines logit-offset magnitudes online at each layer from the reference-attention mass measured during the forward pass. Positive offsets to reference logits strengthen reference usage by edit-region queries, while negative offsets for keep-region queries limit reference-induced changes outside the edit. Two global coefficients control the correction; they are selected once on validation data from four development diffusion editors and held fixed. Across seven diffusion-based image/video editors, RefGAP improves identity fidelity in head swapping and face swapping. RefGAP achieves a fidelity-preservation trade-off comparable to separately tuned constant edit-side biases, without per-approach strength sweeps. Additional experiments on virtual try-on and background replacement evaluate transfer beyond identity editing.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Diffusion-based visual editors sometimes have trouble closely following the details of user-provided reference images when making edits. The authors found that this is because the editing process often pays very little attention to the reference images. They created RefGAP, a way to adjust how much attention is given to references during editing without extra training. This adjustment helps the editor better copy important details from the reference, improving results like face or head swapping in images and videos.
Open → 2609.35708v1