When the Edit Changes the Patient: Measuring Identity Preservation in Counterfactual Retinal Images
2026-08-24 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors studied ways to change medical images to show 'what if' scenarios while keeping the patient's identity the same. They tested three types of text-guided editing methods on eye scan images and found that while all methods could successfully edit images, they differed a lot in keeping the patient's identity. Some methods changed the identity more often, but one called paired-training worked best to keep it intact. The authors suggest that future work should always check identity preservation along with how real and accurate the edited images look.
counterfactual image generationidentity preservationmedical imagingretinal OCTimage editing methodssource-anchored editingstructured-prompt editingpaired-trainingembedding alignmentreferee classifiers
Authors
Andrea Posada, Wenke Karbole, Bach Ngoc Doan, Alexander Weers, Solmaz Abdolrahimzadeh, Maria Patsiamanidi, Kahkashan Haider, Vaishali Khare, Daniel Rueckert, Andrew Lotery, Sobha Sivaprasad, Martin J. Menten
Abstract
Counterfactual medical image generation aims to modify an existing image to reflect a hypothetical scenario in which certain characteristics of the imaged subject are altered, while keeping their identity fixed. Most existing works repurpose established image editing methods, which do not directly supervise identity preservation. Instead, they assume that identity is implicitly preserved by anchoring generation to the source image. This assumption is rarely tested and may fail in domains where biometric cues are subtle, such as retinal optical coherence tomography (OCT). In this work, we explicitly measure identity preservation for three groups of text-conditioned editing methods - source-anchored, structured-prompt, and paired-training - using referee classifiers, embedding alignment scores, and a blind reader study. We find that all methods produce high-quality OCT images with comparable editing success, yet their identity preservation differs markedly. Source-anchored editing frequently alters the depicted subject, while paired-training preserves it best. We argue that future work on medical counterfactual generation must explicitly measure and report identity preservation alongside image realism and editing success.