Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging

2026-07-10Artificial Intelligence

Artificial IntelligenceComputer Vision and Pattern RecognitionMultimedia
AI summary

The authors address a problem where medical imaging models often lack important patient and image information needed to check for performance issues in specific groups. They introduce CAPRA, a method that predicts hidden characteristics from images and uses a small labeled set to calibrate these predictions, helping to find groups with different performance even when metadata is missing. CAPRA works across different types of medical images and helps identify problems that other methods may overlook. It also provides a standardized way to analyze and improve model performance during use without needing extra labels at deployment.

medical imagingmetadatasubgroup analysismodel calibrationpatient-level cross-fittingdataset shiftrobust learningsemantic axesdeployment-time analysislatent slice
Authors
Yawen Li, Yan Li, Zhe Xue, Yingxia Shao, Meiyu Liang, Guanhua Ye
Abstract
Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and many robust-learning methods lose the group structure they rely on. We present CAPRA, a calibrated proxy-axis framework for hidden subgroup analysis under missing metadata. CAPRA predicts image-derived semantic axes, calibrates axis posteriors on a small metadata-labeled split via patient-level cross-fitting, and organizes those posteriors into a calibrated subgroup interface that supports both deployment-time failure analysis and downstream robust learning without requiring subgroup labels at deployment. Across fundus, dermoscopy, and chest radiography, CAPRA reveals disparity patterns missed by metadata-only slicing, remains informative under dataset shift, and produces subgroup partitions that align more closely with explicit failure axes than image-only or latent-slice baselines. The same interface can also be reused by downstream robust learners, although those gains are domain-dependent. Overall, CAPRA turns hidden subgroup analysis under missing metadata into a calibrated, interpretable, and reusable subgroup interface for deployment-time analysis and robust transfer.