Foundation models show mixed reliability on retinal image tasks
FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Retinal image analysis helps detect eye diseases but AI models often struggle when tested on different patient groups or types of images. The authors created a benchmark called FOCUS that tests many AI models across multiple eye disease datasets from different places and imaging conditions. Their results show no single type of model works best everywhere; some models perform well generally, while others need fine-tuning to improve. This work helps to better understand how reliable and fair AI models are for eye disease diagnosis in diverse real-world settings.
What this means in practice
- •For clinical ai developers: Evaluate and compare retinal disease detection models for robustness and fairness across diverse datasets using the FOCUS benchmark.
- •For medical device companies: Improve retinal image analysis tools by testing model generalization and calibration with FOCUS before clinical deployment.$Commercial implications: Enables development of safer, more reliable retinal diagnostic products targeting healthcare providers with multi-dataset validation.
Authors
David Restrepo, Chenwei Wu, Luis Filipe Nakayama, Miguel L. Martins, Stergios Christodoulidis, Maria Vakalopoulou, Enzo Ferrante
Abstract
Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards