Assessing adherence to TRIPOD+AI guidelines in machine learning models for predicting small for gestational age and

Giulia Zamagni1, Camilla Fregona2, Moira Barbieri3

  • 1University of Trieste, Trieste, Italy (Zamagni); Clinical Epidemiology and Public Health Research Unit, Institute for Maternal and Child Health - IRCCS "Burlo Garofolo", Trieste, Italy (Zamagni and Monasta).

Insights

Machine learning models show promise for predicting fetal growth restriction (FGR) and small for gestational age (SGA). However, inconsistent definitions, small sample sizes, and poor reporting limit their reliability and clinical use.

Area of Science:

  • Medical Informatics
  • Perinatal Medicine
  • Machine Learning in Healthcare

Background:

  • Fetal growth restriction (FGR) and small for gestational age (SGA) are critical indicators of perinatal health, impacting long-term outcomes.
  • Accurate prediction of FGR/SGA is essential for timely intervention and improved neonatal care.
  • Machine learning (ML) offers potential for enhancing FGR/SGA prediction using clinical data.

Purpose of the Study:

  • To systematically review machine learning (ML) applications for predicting FGR/SGA.
  • To evaluate the methodological rigor and reporting quality of existing ML models in this field.
  • To identify limitations and areas for improvement in ML-based FGR/SGA prediction.

Main Methods:

  • Systematic literature search conducted in MEDLINE and Scopus following PRISMA 2020 guidelines.
  • Inclusion of studies using ML for FGR/SGA prediction with clinical variables and reporting AUROC or accuracy.
  • Assessment of risk of bias (PROBAST) and adherence to TRIPOD+AI guidelines, including sample size adequacy.

Main Results:

  • 20 studies met inclusion criteria from 272 identified; definitions of FGR/SGA were inconsistent.
  • Variable adherence to TRIPOD+AI guidelines; no models reported fairness or heterogeneity; only 15% reported calibration.
  • Only 30% of studies had adequate sample sizes, suggesting potential overfitting and limited generalizability.

Conclusions:

  • ML models hold potential for FGR/SGA prediction but face significant limitations.
  • Inconsistent definitions, underpowered studies, and poor reporting of calibration hinder clinical translation.
  • Future research requires standardized definitions, robust sample sizes, and comprehensive reporting for reliable ML models.
Abstract