Related Experiment Videos
Retinal image analysis for clinical translation: From deep learning to foundation models, generalization, and
Jiayao Chen1, Xiaozhou Feng2, Hao Hu1
1School of Electronic Information Engineering, Xi'an Technological University, Xi'an 710021, Shaanxi, China.
Abstract:
This review examines retinal image analysis from the perspective of clinical translation rather than benchmark-oriented model comparison. Instead of organizing prior work solely by disease category or model family, we synthesize recent advances through four connected dimensions: imaging modality, task taxonomy, methodological paradigm, and translational bottleneck. We compare major tasks, including classification, detection, segmentation, grading, progression prediction, and treatment-response assessment, across fundus photography, OCT/OCTA, and angiographic imaging. We further review the roles of CNNs, U-Net variants, 3D models, Transformers, graph-based methods, hybrid architectures, foundation models, self-supervised pretraining, and multimodal learning under different data conditions and clinical constraints. Beyond technical progress, we analyze why many high-performing systems still fail to translate reliably into practice, highlighting challenges related to distribution shift, label inconsistency, limited external validation, image-quality control, calibration, fairness, and workflow integration. We argue that the next stage of retinal AI requires transferable representations, robust external evaluation, trustworthy uncertainty handling, and demonstrated value within real-world care pathways, rather than incremental architectural novelty alone.