Prediction of early childhood obesity with machine learning and electronic health record data

Xueqin Pang1, Christopher B Forrest2, Félice Lê-Scherban3

  • 1Department of Biomedical and Health Informatics, Children's Hospital of Philadelphia, Philadelphia, USA.

Insights

Machine learning models predict childhood obesity using electronic health records. XGBoost demonstrated superior performance, outperforming other models in predicting obesity up to age seven.

Area of Science:

  • Pediatric Health
  • Machine Learning
  • Public Health

Background:

  • Childhood obesity is a significant public health concern with long-term health implications.
  • Early identification and intervention are crucial for managing childhood obesity.
  • Electronic Healthcare Record (EHR) data offers a rich resource for developing predictive models.

Purpose of the Study:

  • To develop and compare seven machine learning models for predicting childhood obesity.
  • To utilize EHR data up to age two years for predicting obesity incidence by age seven.
  • To identify the most effective machine learning model for early childhood obesity prediction.

Main Methods:

  • Utilized EHR data from 27,203 pediatric patients.
  • Developed seven distinct machine learning models to predict obesity incidence (BMI > 95th percentile).
  • Evaluated model performance using standard classifier metrics and statistical comparison tests.

Main Results:

  • The XGBoost model achieved the highest performance with an AUC of 0.81.
  • XGBoost significantly outperformed other models in precision, F1-score, accuracy, and specificity.
  • Sensitivity was maintained at 80% across models for fair comparison.

Conclusions:

  • Machine learning models can effectively predict childhood obesity using early EHR data.
  • The XGBoost model shows significant promise for early identification of at-risk children.
  • The developed workflow is adaptable for creating other clinical prediction models from EHR data.
Abstract