Related Experiment Video
Updated: Jan 18, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Development and external validation of an interpretable machine learning-based model for obesity risk prediction in
Mei Xue1,2, Shufang Liu2, Xiaoqian Zhang3
1Graduate School, Beijing University of Chinese Medicine, Beijing, China.
Insights
Machine learning, specifically the XGBoost model, accurately predicts childhood obesity risk using key factors like birth measurements and parental BMI. An online tool translates this model for clinical use.
Area of Science:
- Pediatrics
- Public Health
- Data Science
Background:
- Childhood obesity is a global health issue with complex, not fully understood, causes.
- Predicting obesity risk in children and adolescents is crucial for early intervention.
- Existing predictive models require further validation and interpretation.
Purpose of the Study:
- To develop and validate machine learning models for predicting childhood obesity risk.
- To compare the performance of XGBoost, random forest, light gradient boosting machine, and logistic regression.
- To interpret the best-performing model and create a clinical decision-support tool.
Main Methods:
- Utilized data from 19,024 children (training/testing) and 2,410 (external validation) in Beijing and Tangshan.
- Developed four predictive models: XGBoost, random forest, light gradient boosting machine, and logistic regression.
- Employed SHapley Additive exPlanations (SHAP) for model interpretation and feature selection, creating an online risk assessment tool.
Main Results:
- The XGBoost model achieved superior predictive performance with an AUROC of 0.875 on the external validation set.
- SHAP analysis identified nine key predictors: birth length, parental BMI, sleep duration, physical activity, birth weight, maternal age, delivery mode, and gestational age.
- An online tool was developed, providing individualized risk probabilities and SHAP-based explanations.
Conclusions:
- XGBoost is a highly effective ensemble learning method for predicting childhood obesity.
- The developed digital tool aids clinicians in assessing individual childhood obesity risk.
- Interpretable AI models can enhance clinical decision-making for public health challenges.
Background:
The multifactorial mechanisms driving childhood obesity, a global public health challenge, are yet to be fully elucidated. We aimed to develop and externally validate three widely applied machine learning models alongside logistic regression in 2-18-year-old children and adolescents in Beijing and Tangshan to predict obesity risk. As a further step, we wanted to interpret the optimised model and translate it into a web-based tool to inform clinical decision-making.
Methods:
We analysed data of 19 024 (training/testing) and 2410 (external validation) children and adolescents from Beijing and Tangshan, respectively. Using a set of factors including demographic, familial, socioeconomic, lifestyle, and perinatal variables, we developed four models (light gradient boosting machine, random forest, eXtreme gradient boosting (XGBoost), and logistic regression) and compared their predictive performance. After validation, we selected an optimised model and interpreted it using SHapley Additive exPlanations (SHAP) analysis. Then, we developed an online calculator with interpretable visualisations to enable real-time risk assessment.
Results:
The XGBoost model exhibited superior performance, with an area under the receiver operating characteristic curve (AUROC) of 0.875 on the external validation set, significantly outperforming the logistic regression model (AUROC = 0.718). To identify the minimal feature subset that maintained model efficacy, we incrementally incorporated predictors in the descending order of SHAP importance values while assessing key performance metrics (accuracy, AUROC, and F-beta score). This SHAP-based analysis identified nine key predictors of childhood obesity: birth length, paternal body mass index (BMI), maternal BMI, sleep duration, physical activity, birth weight, maternal age at delivery, delivery mode, and gestational age. The deployed online tool provides individualised risk probabilities and SHAP-derived explanations.
Conclusions:
The XGBoost model in our study was the superior ensemble learning method for predicting childhood obesity. The digital tool integrates this model and can help clinical practitioners determine individuals' risk of childhood obesity.

