Related Experiment Video
Updated: Jan 18, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
NLP-ROPCare: predicting retinopathy of prematurity with admission notes using natural language processing
Yulin Zhang1,2, Shuai Zhao3, Jianbing Ren4
1Shenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Objectives:
Retinopathy of prematurity (ROP) is a leading cause of blindness in children worldwide, requiring more efficient models to help predict treatment-requiring ROP. Our study aimed to develop a new prediction model for ROP occurrence and severity, named NLP-ROPCare, using natural language processing (NLP).
Methods And Analysis:
A retrospective observational study. Infants with a gestational age ≤32 weeks or birth weight ≤2000 g were collected in Guangdong Women and Children Hospital from 2013 to 2022, including 3922 preterm infants with 1106 patients with ROP. Four pretrained language models - BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (Robustly Optimized BERT pretraining Approach), MC-BERT (language pre-training via a Meta Controller) and NEZHA (NEural contextualiZed representation for CHinese lAnguage understanding) - were used for development of NLP prediction models based on free-form texts in the admission notes. For comparison, two machine learning methods (Random Forest and Support Vector Machine) were used to construct prediction models based on 20 structured characteristics previously extracted from the admission notes. Performance evaluating metrics included accuracy, precision, recall, F1 score and area under the curve (AUC).
Results:
The NLP prediction models for ROP occurrence outperformed those for severity. The NEZHA model demonstrated the highest accuracy in predicting ROP occurrence, achieving an F1 score of 89.35% and an AUC of 0.90. Its performance was also better than two machine learning models whose highest F1 was 78% with an AUC equal to 0.87. In addition, the F1 score of RoBERTa (78.44%) was slightly higher than that of NEZHA (77.81%) for predicting ROP severity, and the AUC of RoBERTa also achieved the highest 0.91.
Conclusion:
The NLP-ROPCare combines language models NEZHA and RoBERTa to enable early prediction of ROP occurrence and severity based on unstructured free-form texts in the admission notes of preterm infants, highlighting its value in early prevention of ROP. Further external validation should be carried out to better adjust the model.
More Related Videos
10:32Monitoring Dynamic Growth of Retinal Vessels in Oxygen-Induced Retinopathy Mouse Model
Published on: April 2, 2021
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018