CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration
2026-08-17 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionMachine Learning
AI summaryⓘ
The authors created CoM³eT, a new AI model that combines medical images from pathology and radiology, handling both simple (like labeling) and complex (like outlining) tasks across different image types. It uses a special technique called attention to understand details in various dimensions. CoM³eT performed better than other models in multiple medical image tests and also made generating medical reports easier. Importantly, the authors showed it can be fine-tuned efficiently using less computing power, making it accessible for more researchers and hospitals, even through secure collaboration without sharing patient data directly.
medical foundation modelspathologyradiologysparse and dense predictionmultidimensional contextattention mechanismfine-tuningfederated learningtomographic imagingreport generation
Authors
J. Raphael Schäfer, Kai Geissler, Till Nicke, Chiara Tappermann, Karoline Heber, Eike Petersen, Habib Mergan, Lars Ole Schwen, Nick Weiss, Annika Gerken, Jan Hendrik Moltz, Tom Bisson, Isil Dogan O, Tim-Rasmus Kiehl, Norman Zerbe, Sefer Elezkurtaj, Robin S. Mayer, Nadine Flinner, Peter Wild, Isabel Dahm, Felix Peisen, Heinrich von Busch, Robert Grimm, Sebastian Arndt, Lisa Siegler, Matthias Stefan May, Antje Prasse, Natalia Artysh, Fabian Kiessling, Johannes Lotz
Abstract
Medical foundation models improve generalization when training AI models with limited labeled data, but remain confined to a single specialty, such as pathology or radiology, and to either sparse or dense outputs, such as classification or segmentation. Here, we present CoM$^3$eT (Co-representation Multidimensional Multitask Medical Transformer), a medical vision foundation model that unifies pathology and radiology, sparse and dense predictions, and two- and higher-dimensional inputs by modeling multidimensional context with attention. CoM$^3$eT outperformed other medical foundation models in an open competition spanning five tomographic, four whole-specimen, and three two-dimensional datasets, covering sparse and dense prediction tasks as well as report generation. When adapted across diverse clinical applications, training fewer than 2.5% of parameters achieved performance comparable to full fine-tuning, enabling research without access to high-performance GPU clusters. Applied to federated learning across hospitals, this approach achieved performance comparable to pooled-data training over internet connections and with consumer-grade hardware.