Related Experiment Videos
Machine Learning-Based First-Trimester Antenatal Risk Prediction for Adverse Maternal and Neonatal Outcomes:
Sarah Li1, David Y Y Tan2, Jingxian Zhang2
1Department of Obstetrics and Gynaecology, National University Hospital, Singapore, Singapore, Singapore.
Background:
Maternal outcomes remain inequitable worldwide. Severe morbidity persists, and current risk assessment tools are largely arbitrary, focusing on biomedical factors while overlooking social determinants of health. There is a need for data-driven AI models to improve early pregnancy risk identification and management.
Objective:
The study aimed to develop and internally validate first-trimester AI-based antenatal risk assessment models across three geographically and socioethnically diverse populations (Sweden, Chile, and Singapore) and to compare their performance with existing clinical risk assessment strategies.
Methods:
We conducted a retrospective population-based study using routinely collected first-trimester data from over 700,000 pregnancies from Sweden, Chile, and Singapore. Separate machine learning models predicting a composite of adverse maternal and neonatal outcomes were trained and internally validated for each population. Input variables were limited to information available at or before 14 weeks' gestation. Model discrimination, measured by the area under the receiver operating characteristic (AUROC) curve, was compared with corresponding proxies for real-world first-trimester risk assessment approaches in each setting. Model interpretability was assessed using Shapley additive explanations.
Results:
The prevalence of the composite adverse outcome was 10.40% (75,647/727,354) in Sweden, 21.94% (1302/5934) in Chile, and 16.25% (6145/37,813) in Singapore. In Sweden, the guideline-based risk assessment achieved an AUROC of 0.53, compared with 0.65 for the LightGBM (light gradient boosting machine) model (P<.001). In Chile, the midwifery-led risk assessment achieved an AUROC of 0.52, versus 0.65 from the traditional machine learning LightGBM model (P<.001). In Singapore, the health care professional-based risk assessment reached an AUROC of 0.56, compared with 0.60 for the LightGBM model (P<.05). In the Swedish and Singapore cohorts, sociodemographic variables were among the most influential predictive features. At a false-positive rate of 50.0%, the sensitivities for predicting the primary composite adverse outcome were 69.17%, 69.62%, and 64.93% for the Sweden, Chile, and Singapore models, respectively. At a 30.0% false-positive rate operating point, the sensitivities were lower, at 50.38%, 55.38%, and 44.75%, respectively, but with higher positive predictive values of 16.31%, 34.53%, and 22.32%, respectively. The model calibration plots showed reasonable agreement in Sweden and Singapore, whereas the Chile model showed poorer calibration, with a calibration slope of approximately 1.45, indicating underconfidence.
Conclusions:
AI-based models developed using first-trimester data generally demonstrated improved performance compared with existing first-trimester clinical risk stratification strategies across three distinct populations. These findings suggest the potential feasibility of population-specific, AI-enabled risk stratification as a clinical decision support tool and highlight the potential value of integrating social, demographic, and behavioral determinants into antenatal risk assessment frameworks to support more equitable and personalized antenatal care.