Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

2026-07-27Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors review how color fundus photography (CFP), a way to take pictures inside the eye, has been improved with artificial intelligence (AI). They explain that datasets used for AI have grown bigger and more diverse, while preprocessing techniques have moved from simple fixes to advanced methods that handle missing medical data. AI models themselves have evolved from basic neural networks to more complex systems that combine different types of patient information. The authors suggest future success will need all these parts—data, processing, and models—to work well together for better eye disease detection and clinical use.

Color Fundus PhotographyArtificial IntelligenceDatasetsPreprocessingConvolutional Neural NetworksVision Foundation ModelsMultimodal ModelingElectronic Health RecordsState Space ModelsClinical Screening
Authors
Yu Li, Wengan He, Wenhui Xu, Lihong Jiang, Fan Xiao, Zhuohang Huang, Yuanzhu Liang, Jiayi Liu, Yuxi Chen, Yongsheng Luo
Abstract
Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence. This review provides an integrated overview of CFP AI through the interplay of dataset evolution, preprocessing paradigms, and modeling frameworks. We show that CFP datasets have evolved from small single-center collections with task-specific labels to large multi-center resources featuring multimodal pairings and longitudinal clinical records. Preprocessing has progressed from conventional image enhancement to neural data-engineering pipelines, hardware-aware token optimization, and self-supervised imputation for incomplete electronic health records (EHRs). Meanwhile, modeling has advanced from convolutional neural networks (CNNs) to vision foundation models, state space models (SSMs), and multimodal expert architectures. At the multimodal frontier, CFP is increasingly integrated with EHRs and longitudinal patient information, enabling more comprehensive clinical reasoning beyond isolated image analysis. We conclude that future progress depends on the collaborative optimization of datasets, preprocessing, and multimodal modeling, providing a roadmap toward robust clinical deployment, improved cross-domain generalization, and resource-efficient edge intelligence.