Related Experiment Video
Updated: Sep 18, 2025

Multimodal Volumetric Retinal Imaging by Oblique Scanning Laser Ophthalmoscopy oSLO and Optical Coherence Tomography OCT
Published on: August 4, 2018
A multimodal visual-language foundation model for computational ophthalmology
Danli Shi1,2, Weiyi Zhang3, Jiancheng Yang4
1School of Optometry, The Hong Kong Polytechnic University, Kowloon, Hong Kong SAR, China. danli.shi@polyu.edu.hk.
Abstract:
Early detection of eye diseases is vital for preventing vision loss. Existing ophthalmic artificial intelligence models focus on single modalities, overlooking multi-view information and struggling with rare diseases due to long-tail distributions. We propose EyeCLIP, a multimodal visual-language foundation model trained on 2.77 million ophthalmology images from 11 modalities with partial clinical text. Our novel pretraining strategy combines self-supervised reconstruction, multimodal image contrastive learning, and image-text contrastive learning to capture shared representations across modalities. EyeCLIP demonstrates robust performance across 14 benchmark datasets, excelling in disease classification, visual question answering, and cross-modal retrieval. It also exhibits strong few-shot and zero-shot capabilities, enabling accurate predictions in real-world, long-tail scenarios. EyeCLIP offers significant potential for detecting both ocular and systemic diseases, and bridging gaps in real-world clinical applications.
Related Concept Videos
Visual System
Once through the pupil, the light passes through the lens, a...
Glaucoma: Overview
Vision
Imaging Biological Samples with Optical Microscopy
In optical microscopy, the specimen to be viewed is placed on a glass slide and clipped on the stage...

