Multimodal AI predicts bladder cancer subtypes and progression risks

CHIMERA Challenge Task 2 and 3: Response Subtypes Classification and Progression Survival Prediction in Bladder Cancer Patients using Multimodal Datasets

Computer Vision and Pattern Recognition

Summary

Bladder cancer can come back or get worse, especially in high-risk cases. The authors set up a challenge called CHIMERA to see how well AI can predict specific cancer types and how quickly the disease might progress using tumor images, patient data, and gene activity. They tested many AI models on a large set of patient data and found that combining different data types helps predictions, though some details still cause errors. Their work shows that it’s important for AI methods to handle missing information and to be tested in different hospitals.

What this means in practice

  • For hospital data teams: Integrate multimodal AI models to identify bladder cancer response subtypes and estimate progression risk using combined clinical, image, and gene data.
  • For clinical trial coordinators: Use AI predictions from multimodal data to better stratify high-risk bladder cancer patients for tailored treatment protocols and trial enrollment.

Authors

Catherine Chia, Tongjie Wang, Robert Spaans, Maryam Mohammadlou, Farbod Khoraminia, J. Alberto Nakauma-González, Adam Kowalewski, Parandzem Khachatryan, Domingos Oliveira, Khrystyna Faryna, CHIMERA Challenge Consortium, Marlies Wakkee, Sita Vermeulen, Tahlita Zuiverloon, Nadieh Khalili

Abstract

High-risk non-muscle-invasive bladder cancer (HR-NMIBC) carries substantial risks of recurrence and progression, while current clinical risk stratification remains limited. CHIMERA was established as a multimodal AI challenge to benchmark prediction in HR-NMIBC under standardized evaluation. Task BRS predicts RNA-seq-defined BCG Response Subtypes from histopathology and structured clinicopathological data, whereas Task Progression models time-to-progression using histopathology, structured data, and RNA sequencing. A multimodal dataset of 368 patients was divided into public training and hidden validation and test sets. In total, 159 submissions were made, and 13 top-performing models were selected for benchmarking. The best models achieved a weighted F1 score of 0.73 for Task BRS and a C-index of 0.68 for Task Progression. Post-challenge analyses revealed task-dependent modality contributions, cohort-dependent performance degradation, and sensitivity to missing structured data. In Task BRS, histopathology partly compensated for pathology-derived structured variables, whereas progression models showed greater dependence on complementary inputs. Cross-model error analysis further identified patients that were consistently difficult across different architectures, with T1 substage associated with prediction difficulty. These findings highlight barriers to transportability and the importance of missingness-aware modeling and independent multi-institutional validation. CHIMERA provides a standardized multimodal benchmark for bladder cancer and a framework for studying not only model performance, but also robustness, information sufficiency, and patient-level prediction failure.