Related Experiment Video
Updated: Aug 22, 2025

Author Spotlight: Assessing the Feasibility of Using Amplitude-Integrated EEG During Neonatal Transport
Published on: June 21, 2024
Development of prognostic model for preterm birth using machine learning in a population-based cohort of Western
Kingsley Wong1,2, Gizachew A Tessema3,4, Kevin Chai3
1Curtin School of Population Health, Curtin University, 400 Kent St, Bentley, Perth, WA, 6102, Australia. kingsley.wong@postgrad.curtin.edu.au.
Insights
Predicting preterm birth is crucial for public health. Machine learning models using routine maternal data can identify nearly half of all preterm births antenatally with high specificity.
Area of Science:
- Obstetrics and Gynecology
- Public Health
- Medical Informatics
Background:
- Preterm birth presents a significant global health challenge.
- Existing prognostic models for preterm birth require enhancement.
- Population-based data offers a valuable resource for predictive modeling.
Purpose of the Study:
- To develop and validate machine learning models for preterm birth prediction.
- To utilize routinely collected, population-based data for model development.
- To assess the performance of various classification algorithms in predicting preterm birth.
Main Methods:
- A longitudinal retrospective cohort study of births in Western Australia (1980-2015).
- Development of prediction models using logistic regression, decision trees, Random Forests, extreme gradient boosting, and multi-layer perceptron (MLP).
- Inclusion of maternal socio-demographics, medical conditions, pregnancy complications, and family history as predictors; stratified tenfold cross-validation was employed.
Main Results:
- The best performing model (MLP) correctly classified 49.1% of preterm births at 5% false positive rate using current pregnancy data.
- Including past obstetric history improved sensitivity to 52.7% in multiparous women.
- Approximately half of preterm births can be identified antenatally with high specificity.
Conclusions:
- Machine learning models can effectively predict preterm birth using routinely collected maternal and pregnancy data.
- Model performance is influenced by the availability and type of predictor variables.
- Antenatal identification of nearly half of preterm births is achievable, aiding targeted interventions.
Abstract:
Preterm birth is a global public health problem with a significant burden on the individuals affected. The study aimed to extend current research on preterm birth prognostic model development by developing and internally validating models using machine learning classification algorithms and population-based routinely collected data in Western Australia. The longitudinal retrospective cohort study involved all births in Western Australia between 1980 and 2015, and the analytic sample contains 81,974 (8.6%) preterm births (< 37 weeks of gestation). Prediction models for preterm birth were developed using regularised logistic regression, decision trees, Random Forests, extreme gradient boosting, and multi-layer perceptron (MLP). Predictors included maternal socio-demographics and medical conditions, current and past pregnancy complications, and family history. Class weight was applied to handle imbalanced outcomes and stratified tenfold cross-validation was used to reduce overfitting. Close to half of the preterm births (49.1% at 5% FPR, 95% CI 48.9%,49.5%) were correctly classified by the best performing classifier (MLP) for all women when current pregnancy information was available. The sensitivity was boosted to 52.7% (95% CI 52.1%,53.3%) after including past obstetric history in a sub-population of births from multiparous women. Around half of the preterm birth can be identified antenatally at high specificity using population-based routinely collected maternal and pregnancy data. The performance of the prediction models depends on the available predictor pool that is individual and time specific.
More Related Videos
Related Concept Videos
Steps in Outbreak Investigation
Regression Toward the Mean

