The pediatric transfer gap: How adult-ICU-trained machine-learning models degrade on pediatric intensive care

Ji-Young Yeo1, Eun Sun So2, Sungkwan Youm3

  • 1AI Institute, Hanyang University, 222, Wangsimni-ro, Seongdong-gu, Seoul, 04763, Republic of Korea.

Insights

Clinical machine-learning models trained on adult data perform poorly in children, especially neonates. Limited pediatric data and retraining can significantly improve model accuracy and calibration for pediatric populations.

Area of Science:

  • Pediatric critical care medicine
  • Machine learning in healthcare
  • Clinical data science

Background:

  • Machine learning models trained on adult data pose safety risks when applied to pediatric populations.
  • The "Pediatric Transfer Gap" quantifies performance degradation in pediatric models due to adult data training.
  • This study evaluates the impact of the Pediatric Transfer Gap and the efficacy of limited target-cohort data in mitigating it.

Purpose of the Study:

  • To quantify the Pediatric Transfer Gap in routinely recorded vital signs for pediatric patients.
  • To assess whether limited pediatric target-cohort data can effectively close this gap.
  • To evaluate the performance and calibration of machine learning models across different pediatric age bands.

Main Methods:

  • Two experiments were conducted: a within-pediatric design (Experiment A) and an adult-ICU to pediatric transfer design (Experiment B).
  • Experiment A used data from 10,194 pediatric intensive care patients (older children as source, younger as target).
  • Experiment B utilized data from 63,203 adult ICU patients, transferring their vital signs to a pediatric cohort.

Main Results:

  • Models trained on older children showed a significant drop in ROC-AUC (0.940 to 0.725) when applied to younger children, with retraining recovering performance (0.867).
  • Domain shift was evident in Experiment B, primarily as severe miscalibration (ECE rising to 0.285), which was resolved by retraining.
  • Model performance decreased with age, with neonates showing the lowest performance (ROC-AUC 0.587) and highest miscalibration (ECE 0.249).

Conclusions:

  • Vital-sign-based mortality models trained on adult or older pediatric data lose discrimination and calibration when applied to younger children, particularly neonates.
  • Target-domain retraining, even with a small amount of data (5%), significantly improves model performance and calibration.
  • Pediatric-specific validation, calibration-aware reporting, and retraining are crucial before clinical implementation of these models.
Abstract

Related Concept Videos