Development and validation of machine learning classifiers for predicting treatment-needed retinopathy of prematurity
Nasser Shoeibi1, Majid Abrishami1, Seyedeh Maryam Hosseini1
1Eye Research Center, Mashhad University of Medical Sciences, Mashhad, Iran.
Insights
Machine learning models can identify premature infants needing treatment for retinopathy of prematurity (ROP). The Naïve Bayes model showed the highest sensitivity, crucial for clinical decisions in neonates.
Area of Science:
- Medical Informatics
- Neonatology
- Machine Learning in Healthcare
Background:
- Retinopathy of prematurity (ROP) is a significant concern in premature infants.
- Accurate identification of neonates requiring ROP treatment is critical for timely intervention.
Purpose of the Study:
- To design and evaluate supervised machine learning models for ROP treatment identification.
- To assess model performance using demographic and clinical data from screened premature infants.
Main Methods:
- Retrospective review of 9,692 infants screened for ROP.
- Extraction of eleven demographic and clinical features.
- Development and assessment of eight machine learning classifiers (LR, DT, SVM, NB, KNN, XGBoost, ANN, RF).
Main Results:
- XGBoost and Artificial Neural Networks (ANN) achieved 96% accuracy.
- Naïve Bayes (NB) demonstrated the highest sensitivity (0.99), indicating the lowest false negative rate.
- A model with high sensitivity is clinically preferable for identifying neonates requiring ROP treatment.
Conclusions:
- AI tools can augment clinical decision-making in ROP management.
- Model outputs should supplement, not replace, clinical judgment.
- Clinicians must integrate AI insights with holistic patient assessment for final treatment decisions.
Background:
This study aims to design and evaluate various supervised machine-learning models for identifying premature infants who require treatment based on demographic data and clinical findings from screening examinations.
Methods:
We conducted a retrospective review of medical records for infants screened for retinopathy of prematurity (ROP) at our clinic over the past decade. We extracted demographic and clinical data, including eleven features: sex, maternal education, paternal education, birth weight, gestational age, ROP stage, zone of retinal involvement, age at examination, weight at examination, and CPR. We developed and assessed several classifiers: logistic regression (LR), decision tree (DT), support vector machine (SVM), naïve Bayes (NB), K-nearest neighbors (KNN), XGBoost, artificial neural networks (ANN), and random forest (RF). The target variable was defined as whether the neonate received any treatment during the follow-up period.
Results:
Our analysis included data from 9,692 infants. Among the machine learning models evaluated, the XGBoost and ANN models achieved the highest accuracy at 96%. In terms of sensitivity (recall), the NB model exhibited the lowest false negative rate, indicating the highest sensitivity (0.99). In the context of premature neonates, accurately diagnosing those who require treatment is crucial. Therefore, from a clinical perspective, prioritizing a model with the lowest false negative rate may be more beneficial than selecting one based solely on the highest accuracy.
Conclusion:
While AI can enhance decision-making processes by providing real-time risk assessments, these tools must be used to augment-not replace-clinical judgment. Clinicians must remain involved in interpreting model outputs and making final treatment decisions based on a holistic understanding of each patient's unique circumstances.
Clinical Trial Number:
Not applicable.


