Related Experiment Video
Updated: Jul 9, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.8K
Validating the Generalizability of Ophthalmic Artificial Intelligence Models on Real-World Clinical Data.
Homa Rashidisabet1,2, Abhishek Sethi2,3, Ponpawee Jindarak3
1Department of Biomedical Engineering, University of Illinois Chicago, Chicago, IL, USA.
Translational Vision Science & Technology
|November 3, 2023
Summary
Deep learning models trained on public fundus images struggle to generalize to real-world data for glaucoma diagnosis. Real-world data is key to improving these models for clinical use.
Area of Science:
- Ophthalmology
- Medical Imaging
- Artificial Intelligence
Background:
- Deep learning (DL) models show promise for diagnosing eye conditions like glaucoma using fundus images.
- However, models trained on public datasets may not perform well on real-world clinical data due to variations in image acquisition and patient populations.
Purpose of the Study:
- To assess the generalizability of deep learning models trained on public fundus image datasets when applied to real-world data for glaucoma diagnosis.
- To compare the performance of models trained on public versus real-world data for glaucoma classification and optic disc segmentation.
Main Methods:
- Utilized six public fundus image datasets and one real-world dataset (Illinois Eye and Ear Infirmary).
- Trained and tested deep learning models for glaucoma classification and optic disc segmentation on both public and real-world datasets.
- Analyzed model decision-making processes and learned embeddings for glaucoma classification.
Main Results:
- Publicly trained models performed better on public test data (95.0% AUC for classification, 96.3% IoU for segmentation).
- Performance of publicly trained models significantly decreased on real-world test data (76.6% AUC, 85.6% IoU).
- Models trained on real-world data showed superior performance on real-world test sets compared to publicly trained models.
Conclusions:
- Deep learning models trained on public fundus images exhibit limited generalizability for glaucoma diagnosis in real-world settings.
- Real-world data is crucial for enhancing the generalizability of deep learning models and facilitating their clinical translation in ophthalmology.

