Related Experiment Video
Updated: Sep 15, 2025

Assessment of Child Anthropometry in a Large Epidemiologic Study
Published on: February 2, 2017
Application and Analysis of Random Forest and Support Vector Classification in Risk Prediction of Childhood Obesity
Yuhang Wang1, Shuang Shi1, Xinghua Wei1
1Graduate School, Nantong University, Nantong, Jiangsu, People's Republic of China.
Insights
Machine learning models effectively predict childhood obesity and hyperuricemia risk. These tools aid early detection and intervention for better long-term health outcomes in children.
Area of Science:
- Pediatric Health
- Metabolic Disorders
- Machine Learning in Medicine
Background:
- Childhood obesity and hyperuricemia are rising public health concerns.
- These conditions increase cardiometabolic disease risk through complex metabolic interactions.
- Machine learning (ML) offers a promising approach for pediatric risk prediction.
Purpose of the Study:
- To develop and evaluate two ML models: Random Forest (RF) and Support Vector Classification (SVC).
- To predict the risk of childhood obesity and hyperuricemia by integrating clinical and biochemical data.
- To assess model performance using AUC, precision-recall, and calibration curves, and interpret feature importance with SHAP analysis.
Main Methods:
- 101 children (60 obese, 41 obese with hyperuricemia) were enrolled.
- Data preprocessing included recursive feature elimination (RFE), ROSE oversampling, and standardization.
- RF and SVC models were trained and validated; SHAP analysis identified key predictors.
Main Results:
- Both RF and SVC models achieved high predictive performance with AUCs of 0.96.
- SVC showed higher precision and recall, suitable for community screening.
- RF demonstrated superior calibration, beneficial for clinical decision-making.
- Key predictors identified include GFR, HDL-C, and ApoB, with some nonlinear associations.
Conclusions:
- RF and SVC models provide reliable tools for early risk prediction of childhood obesity and hyperuricemia.
- Models are tailored for distinct clinical scenarios, supporting early identification and targeted interventions.
- Future research will explore metabolomic data and ensemble methods to enhance performance.
Background:
The concurrent rise of childhood obesity and hyperuricemia presents a serious public health concern. These conditions interact through complex metabolic mechanisms and significantly increase long-term risks of cardiometabolic diseases. Machine learning (ML) offers an effective framework for constructing efficient risk prediction models in pediatric populations.
Objective:
This study aimed to develop and evaluate two ML models-Random Forest (RF) and Support Vector Classification (SVC)-to predict the risk of childhood obesity and hyperuricemia by integrating clinical and biochemical variables.
Methods:
A total of 101 children were enrolled, including 60 with obesity and 41 with obesity plus hyperuricemia. Data preprocessing involved recursive feature elimination (RFE), ROSE-based oversampling, and feature standardization. Both RF and SVC models were trained and evaluated using area under the ROC curve (AUC), precision-recall curves, and calibration curves. SHAP (Shapley Additive Explanations) analysis was conducted to interpret feature contributions.
Results:
Both models demonstrated strong predictive performance, with AUCs reaching 0.96. The SVC model achieved slightly higher average precision and recall, making it more suitable for community- or school-based screening of high-risk children. In contrast, the RF model exhibited superior calibration, suggesting its greater utility in clinical decision-making where probabilistic risk estimation guides personalized follow-up or intervention planning. SHAP analysis identified glomerular filtration rate (GFR), high-density lipoprotein cholesterol (HDL-C), and apolipoprotein B (ApoB) as key predictors, some exhibiting nonlinear associations with disease risk.
Conclusion:
RF and SVC models offer reliable tools for early risk prediction of obesity and hyperuricemia in children, each tailored to distinct clinical scenarios. These findings support early identification and targeted intervention. Future studies will explore the integration of metabolomic data and ensemble approaches to further enhance model performance and clinical applicability.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Obesity
Receiver Operating Characteristic Plot
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Cancer Survival Analysis
Regression Toward the Mean

