Multimodal models improve clinical diagnosis with better ranking methods

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Machine LearningComputation and LanguageComputer Vision and Pattern Recognition

Summary

Standard methods to teach AI models about clinical data often focus on accuracy, which can be misleading when one outcome is much more common than others. The authors show that focusing on a ranking measure called AUROC, which better handles unbalanced data, helps these AI systems make more reliable clinical decisions. They improve the way AI models choose their prompts by comparing pairs of patient cases instead of evaluating each case alone. Their approach helps the AI rank diseases more effectively across several clinical tasks with no extra computation.

What this means in practice

  • For hospital data teams: Improve AI diagnostic tools by optimizing prompts for better disease ranking despite class imbalance in clinical data.
  • For medical ai developers: Design multimodal AI diagnosis systems that incorporate pairwise ranking feedback to enhance clinical decision-making accuracy.

Authors

Tian Xia, Minghao Liu, Yiqing Liang, Laixi Shi, Jiayun Wang

Abstract

Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.