Related Experiment Video
Updated: May 22, 2026

Assessment of Child Anthropometry in a Large Epidemiologic Study
Published on: February 2, 2017
Development and validation of an explainable machine learning-based risk prediction model for obesity in Chinese
Zekai Chen1, Lin Zhu2,3, Peijie Chen1
1School of Exercise and Health, Shanghai University of Sport, Shanghai, China.
Insights
A new machine learning model accurately predicts childhood obesity risk in China. Key factors include parental BMI, screen time, and physical activity, enabling early intervention for this public health challenge.
Area of Science:
- Public Health
- Pediatrics
- Machine Learning
Background:
- Childhood obesity is a major global health concern requiring effective prediction tools.
- Limited interpretable models exist for identifying obesity risk in children and adolescents using national data.
- This study aimed to develop and validate a predictive model for childhood obesity in China.
Purpose of the Study:
- To develop and validate an interpretable and user-friendly obesity risk prediction model for Chinese children and adolescents.
- To utilize nationally representative data for accurate risk assessment.
- To facilitate early prevention and intervention strategies for childhood obesity.
Main Methods:
- Utilized cross-sectional data (2017-2018) and temporal validation (2020) from the Physical Activity and Fitness in China-The Youth Study (PAFCTYS).
- Employed Least Absolute Shrinkage and Selection Operator (LASSO) with recursive feature elimination (RFE) to identify key predictors from 38 candidate variables.
- Developed and compared eight machine learning algorithms, selecting the Random Forest (RF) model for its superior performance and interpreting it using SHapley Additive exPlanation (SHAP).
Main Results:
- The Random Forest (RF) model demonstrated excellent performance with an AUC of 0.946 on the test set and 0.810 on the temporal validation set.
- Key predictors identified by SHAP analysis include parental BMI, weekday mobile device use, moderate-to-vigorous physical activity (MVPA), weekday TV watching, and sex.
- A web-based risk calculator based on the RF model was successfully deployed.
Conclusions:
- Developed and validated an explainable machine learning model for predicting childhood obesity risk in China using a large, national sample.
- The model accurately assesses current obesity risk, aiding healthcare, schools, and parents in large-scale screening.
- The findings support early identification and intervention for childhood obesity.
Background:
Childhood obesity represents a significant global public health challenge. Accurate and rapid prediction models for identifying obesity risk in children and adolescents are essential for facilitating early prevention and enabling timely interventions. However, interpretable and user-friendly obesity risk prediction models based on nationally representative data remain limited. This study aimed to develop and validate a model to predict current obesity risk among children and adolescents in China.
Methods:
The models were developed using cross-sectional data from the 2017-2018 Physical Activity and Fitness in China-The Youth Study (PAFCTYS; n = 35,016) and were temporally validated with 2020 data (n = 3,495). Candidate predictors (n = 38), primarily encompassing physical activity, sedentary behavior, and sociodemographic variables, were measured concurrently with the outcome. The participants included individuals from 31 administrative regions across China. Features were identified using the Least Absolute Shrinkage and Selection Operator (LASSO) with recursive feature elimination (RFE), and predictive models were developed using eight different machine learning algorithms. The model's performance was evaluated using metrics such as AUC, sensitivity, specificity, Matthews Correlation Coefficient (MCC), and Brier score. The model was interpreted using the SHapley Additive exPlanation (SHAP) method.
Results:
The random forest (RF) model outperformed all other models, achieving an AUC of 1.000 on the training set and 0.946 on the testing set. It also maintained good discriminative performance on the temporal validation dataset, with an AUC of 0.810. The RF model also shows the highest accuracy, specificity, and MCC, and the lowest Brier score. SHAP analysis indicates that parental BMI, using mobile electronic devices on weekdays, MVPA (moderate-to-vigorous physical activity), watching TV on weekdays, and sex are the top 5 most important features in the obesity risk prediction model. A web-based risk calculator based on the RF model has been successfully deployed.
Conclusion:
Using a large, nationally representative sample and readily available variables, we successfully developed and validated an obesity risk prediction model for Chinese children and adolescents using explainable machine learning techniques. The model can quickly and accurately assess an individual's current obesity risk, enabling healthcare agencies, schools, and parents to conduct large-scale obesity risk screening.
