Related Experiment Video
Updated: Jul 5, 2025

Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
Image Analysis-Based Machine Learning for the Diagnosis of Retinopathy of Prematurity: A Meta-analysis and Systematic
Yihang Chu1, Shipeng Hu2, Zilan Li3
1Central South University of Forestry and Technology, Changsha, Hunan, China; State Key Laboratory of Pathogenesis, Prevention and Treatment of High Incidence Diseases in Central Asia, Clinical Medical Research Institute, The First Affiliated Hospital of Xinjiang Medical University, Urumqi, Xinjiang, China.
Insights
Machine learning (ML) shows high accuracy in diagnosing retinopathy of prematurity (ROP), comparable to human experts. While promising for automated diagnosis, ML tools currently serve best as supplementary aids for clinicians.
Area of Science:
- Ophthalmology
- Medical Imaging
- Artificial Intelligence
Background:
- Retinopathy of prematurity (ROP) is a significant cause of blindness in preterm infants.
- Early detection and diagnosis of ROP are critical for effective treatment and prevention of vision loss.
Purpose of the Study:
- To systematically evaluate the diagnostic performance of machine learning (ML) algorithms for retinopathy of prematurity (ROP).
- To assess the potential of ML as an automated diagnostic tool in clinical settings for ROP detection and classification.
Main Methods:
- A systematic review and meta-analysis of published studies on image-based ML for ROP diagnosis and subtype classification.
- Searches conducted across major databases (Web of Science, PubMed, Embase, IEEE Xplore, Cochrane Library) up to October 2022.
- Quality assessment using AI-centered diagnostic accuracy tools and statistical analysis including bivariate mixed-effects models and Deek's test.
Main Results:
- Twenty-two studies were included, with varying degrees of risk of bias and applicability concerns.
- Image-based ML demonstrated high diagnostic performance for ROP, with sensitivity of 93% and specificity of 95% (AUC=0.98).
- ML also showed high accuracy in classifying ROP subtypes (sensitivity 93%, specificity 93%, AUC=0.97), comparable to clinical experts (Spearman's R=0.879).
Conclusions:
- Machine learning algorithms exhibit diagnostic accuracy for ROP that is non-inferior to that of human experts.
- ML holds significant potential as an automated tool for ROP diagnosis and classification.
- Due to evidence heterogeneity and quality, ML tools are currently best utilized as supplementary aids to assist clinicians in ROP diagnosis.
Topic:
To evaluate the performance of machine learning (ML) in the diagnosis of retinopathy of prematurity (ROP) and to assess whether it can be an effective automated diagnostic tool for clinical applications.
Clinical Relevance:
Early detection of ROP is crucial for preventing tractional retinal detachment and blindness in preterm infants, which has significant clinical relevance.
Methods:
Web of Science, PubMed, Embase, IEEE Xplore, and Cochrane Library were searched for published studies on image-based ML for diagnosis of ROP or classification of clinical subtypes from inception to October 1, 2022. The quality assessment tool for artificial intelligence-centered diagnostic test accuracy studies was used to determine the risk of bias (RoB) of the included original studies. A bivariate mixed effects model was used for quantitative analysis of the data, and the Deek's test was used for calculating publication bias. Quality of evidence was assessed using Grading of Recommendations Assessment, Development and Evaluation.
Results:
Twenty-two studies were included in the systematic review; 4 studies had high or unclear RoB. In the area of indicator test items, only 2 studies had high or unclear RoB because they did not establish predefined thresholds. In the area of reference standards, 3 studies had high or unclear RoB. Regarding applicability, only 1 study was considered to have high or unclear applicability in terms of patient selection. The sensitivity and specificity of image-based ML for the diagnosis of ROP were 93% (95% confidence interval [CI]: 0.90-0.94) and 95% (95% CI: 0.94-0.97), respectively. The area under the receiver operating characteristic curve (AUC) was 0.98 (95% CI: 0.97-0.99). For the classification of clinical subtypes of ROP, the sensitivity and specificity were 93% (95% CI: 0.89-0.96) and 93% (95% CI: 0.89-0.95), respectively, and the AUC was 0.97 (95% CI: 0.96-0.98). The classification results were highly similar to those of clinical experts (Spearman's R = 0.879).
Conclusions:
Machine learning algorithms are no less accurate than human experts and hold considerable potential as automated diagnostic tools for ROP. However, given the quality and high heterogeneity of the available evidence, these algorithms should be considered as supplementary tools to assist clinicians in diagnosing ROP.
Financial Disclosure(S):
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

