Related Experiment Video
Updated: Oct 21, 2025

One-step Metabolomics: Carbohydrates, Organic and Amino Acids Quantified in a Single Procedure
Published on: June 25, 2010
Machine learning guided postnatal gestational age assessment using new-born screening metabolomic data in South Asia
Sunil Sazawal1, Kelli K Ryckman2, Sayan Das3
1Center for Public Health Kinetics, Global Division, 214 A, LGL Vinoba Puri, Lajpat Nagar II, New Delhi, India. ssazawal@jhu.edu.
Insights
Machine learning accurately estimates gestational age in low-resource settings using metabolomics, improving tracking of early and small births to reduce infant mortality. This approach surpasses traditional methods for preterm birth identification.
Area of Science:
- Computational biology and machine learning applications in global health.
- Neonatal and perinatal epidemiology in low and middle-income countries (LMICs).
Background:
- Premature and small-for-gestational-age births in LMICs significantly contribute to neonatal and infant mortality.
- Current methods for estimating gestational age (GA) in LMICs (new-born assessment, LMP, birth weight) are unreliable due to lack of early ultrasound access.
- Existing algorithms from developed settings use metabolic screen data for GA estimation, showing potential for LMICs.
Purpose of the Study:
- To develop and evaluate machine learning (ML) algorithms for accurate gestational age estimation using metabolomic data in LMICs.
- To improve population-level tracking of early and small for gestational age births for policy and care.
Main Methods:
- Utilized prospective pregnancy cohort data (AMANHI-ACT) from Asia and Africa, including ultrasonography-based GA, birth weight, and newborn metabolomic screening data from 1318 infants.
- Developed and tested Random Forest Regressor ML models, splitting data into training and testing sets.
- Evaluated model performance using Mean Absolute Error (MAE) and Root Mean Square Error (RMSE), with bootstrap for confidence intervals. ROC analysis was used for preterm birth classification.
Main Results:
- The ML model achieved a Mean Absolute Error (MAE) of 5.2 days for gestational age estimation, comparable to performance in small-for-gestational-age (SGA) estimation (MAE 5.3 days).
- Gestational age was accurately estimated within one week for 85.21% of newborns.
- For preterm birth classification, the model demonstrated high accuracy with an Area Under the Curve (AUC) of 98.1%, outperforming the Iowa regression method (AUC difference 14.4%).
Conclusions:
- Machine learning applied to metabolomic data provides a viable method for accurate, population-level gestational age dating in LMICs.
- These findings support the potential for developing region-specific models and exploring focused or broad metabolomic analyses for improved GA estimation.
- This approach offers a significant opportunity to enhance public health strategies and targeted interventions for vulnerable newborns in resource-limited settings.
Background:
Babies born early and/or small for gestational age in Low and Middle-income countries (LMICs) contribute substantially to global neonatal and infant mortality. Tracking this metric is critical at a population level for informed policy, advocacy, resources allocation and program evaluation and at an individual level for targeted care. Early prenatal ultrasound examination is not available in these settings, gestational age (GA) is estimated using new-born assessment, last menstrual period (LMP) recalls and birth weight, which are unreliable. Algorithms in developed settings, using metabolic screen data, provided GA estimates within 1-2 weeks of ultrasonography-based GA. We sought to leverage machine learning algorithms to improve accuracy and applicability of this approach to LMICs settings.
Methods:
This study uses data from AMANHI-ACT, a prospective pregnancy cohorts in Asia and Africa where early pregnancy ultrasonography estimated GA and birth weight are available and metabolite screening data in a subset of 1318 new-borns were also available. We utilized this opportunity to develop machine learning (ML) algorithms. Random Forest Regressor was used where data was randomly split into model-building and model-testing dataset. Mean absolute error (MAE) and root mean square error (RMSE) were used to evaluate performance. Bootstrap procedures were used to estimate confidence intervals (CI) for RMSE and MAE. For pre-term birth identification ROC analysis with bootstrap and exact estimation of CI for area under curve (AUC) were performed.
Results:
Overall model estimated GA had MAE of 5.2 days (95% CI 4.6-6.8), which was similar to performance in SGA, MAE 5.3 days (95% CI 4.6-6.2). GA was correctly estimated to within 1 week for 85.21% (95% CI 72.31-94.65). For preterm birth classification, AUC in ROC analysis was 98.1% (95% CI 96.0-99.0; p < 0.001). This model performed better than Iowa regression, AUC Difference 14.4% (95% CI 5-23.7; p = 0.002).
Conclusions:
Machine learning algorithms and models applied to metabolomic gestational age dating offer a ladder of opportunity for providing accurate population-level gestational age estimates in LMICs settings. These findings also point to an opportunity for investigation of region-specific models, more focused feasible analyte models, and broad untargeted metabolome investigation.

