Federated learning improves missing clinical data filling across centers
Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness
Machine LearningArtificial Intelligence
Summary
When hospitals and research centers collect different types of patient data but want to use it together, sharing raw data is often not allowed and some data features may be entirely missing in some places. The authors created Fed-ReMasker, a tool that helps predict and fill in missing data without sharing sensitive patient information, by learning patterns from all involved centers. Their tests show that this method guesses missing values more accurately than other existing methods, even when some centers have very different data. It works nearly as well as if all data were combined in one place, which isn’t usually possible.
What this means in practice
- •For hospital data teams: Fill in entirely missing patient data features at each hospital while preserving data privacy across collaborating centers.
- •For biomedical data engineers: Improve the accuracy of imputation in federated clinical datasets where centers collect different features under varying protocols.
Authors
Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
Abstract
Multi-center clinical studies and biomedical research collaborations increasingly seek to utilize data across centers to build models that generalize beyond any single center. This creates two distinct challenges: data protection regulations may restrict the sharing of raw patient data across institutions, while centers may collect only partially overlapping sets of features under different protocols. Federated learning enables collaborative model training without centralizing raw data. However, existing federated imputation methods rarely evaluate feature-level missingness, in which entire features are unobserved at some centers. To address this setting, we adapt the ReMasker masked autoencoder to federated learning (Fed-ReMasker), enabling centers to impute features never observed locally by leveraging knowledge learned across collaborating centers. We evaluate Fed-ReMasker in a benchmark spanning synthetic datasets with linear and nonlinear relationships and real-world tabular datasets, including clinical data. The benchmark varies the number of centers, the missingness ratios, and client heterogeneity. Fed-ReMasker achieves the lowest imputation error in 93.2% of value-level and 96.7% of feature-level scenarios in the homogeneous benchmark. It also remains robust to client heterogeneity using simple federated averaging, outperforming all baselines in all 36 value-level scenarios and each baseline in at least 35 of 36 feature-level scenarios, and comes within 3.0% on average of a centralized model trained on the pooled data.