Consistency score measures photographic similarity in surgical photos
Automated Perceptually-Motivated Assessment of Photographic Consistency in Paired Clinical Photographs: Pipeline Development and Internal Evaluation
Computer Vision and Pattern Recognition
Summary
Photos taken before and after plastic surgery need to be consistent to fairly compare results, but there's no standard way to check that. The authors created a system that looks at many details like lighting, sharpness, and angle to give a single score showing how similar two photos are. They tested this score on many photo pairs and found it can reliably tell if photos of the same person were taken under similar conditions. This helps ensure fair comparisons without judging the surgery quality itself.
What this means in practice
- •For plastic surgeons: Check whether pre- and post-operative images were taken under comparable conditions to ensure fair visual outcome assessments.
- •For medical photography teams: Audit and validate clinical photo pairs automatically to improve quality control before documentation or publication.
Authors
Derrick Lin, Samantha Rabinovich, Joclin Rabinovich, Kassra Garoosi, Sumun Khetpal, Evan Delanoy, Neel Bhardwaj, Jason Roostaeian
Abstract
Purpose: Paired pre- and post-operative photographs are the standard unit of evidence for plastic surgical outcomes, yet no objective metric verifies whether two images of the same patient were captured under conditions consistent for comparison. Approach: We developed a perceptually motivated pipeline that analyzes pre/post pairs across thirteen calibrated sub-metrics, partitioned by unsupervised correlation-structure analysis into five data-driven clusters (photometric, texture / sharpness, pose, illumination direction, and pitch), averaged within each cluster and combined across clusters by a weighted sum into a single consistency score. Each sub-metric is calibrated so that its median difference across published within-patient pairs scores 0.5, which is a reference point and carries no pass/fail meaning. The pipeline was calibrated on 134 matched within-patient published pre/post pairs and evaluated against identical-image pairs, synthetic-perturbation pairs, and 134 mismatched cross-publication pairs. Results: The master consistency score S separated matched from mismatched pairs (sensitivity index d' = 2.15, 95% confidence interval (CI) [1.83, 2.55]; area under the receiver operating characteristic curve AUC = 0.928, 95% CI [0.896, 0.959]), closely matching Gaussian-equal-variance predictions. The three head-pose angles did not fall in one cluster: yaw and roll grouped together while pitch separated. Identical pairs scored at ceiling (S = 0.99) and the master score fell monotonically with perturbation magnitude on all five perturbation axes. Conclusions: The score quantifies photographic comparability, not aesthetic or surgical quality, and provides a freely available web tool for auditing the photographic comparability of pre/post pairs, pending validation against expert judgment.