Related Experiment Video
Updated: Jun 12, 2025

Assessment of Child Anthropometry in a Large Epidemiologic Study
Published on: February 2, 2017
Predicting age at onset of childhood obesity using regression, Random Forest, Decision Tree, and K-Nearest
Salem Hamoud Alanazi1,2, Mali Abdollahian1, Laleh Tafakori1
1School of Science, RMIT University, Melbourne, Victoria, Australia.
Abstract:
Childhood and adolescent overweight and obesity are one of the most serious public health challenges of the 21st century. A range of genetic, family, and environmental factors, and health behaviors are associated with childhood obesity. Developing models to predict childhood obesity requires careful examination of how these factors contribute to the emergence of childhood obesity. This paper has employed Multiple Linear Regression (MLR), Random Forest (RF), Decision Tree (DT), and K-Nearest Neighbour (KNN) models to predict the age at the onset of childhood obesity in Saudi Arabia (S.A.) and to identify the significant factors associated with it. De-identified data from Arar and Riyadh regions of S.A. were used to develop the prediction models and to compare their performance using multi-prediction accuracy measures. The average age at the onset of obesity is 10.8 years with no significant difference between boys and girls. The most common age group for onset is (5-15) years. RF model with the R2 = 0.98, the root mean square error = 0.44, and mean absolute error = 0.28 outperformed other models followed by MLR, DT, and KNN. The age at the onset of obesity was linked to several demographic, medical, and lifestyle factors including height and weight, parents' education level and income, consanguineous marriage, family history, autism, gestational age, nutrition in the first 6 months, birth weight, sleep hours, and lack of physical activities. The results can assist in reducing the childhood obesity epidemic in Saudi Arabia by identifying and managing high-risk individuals and providing better preventive care. Furthermore, the study findings can assist in predicting and preventing childhood obesity in other populations.
Related Concept Videos
Regression Toward the Mean
Survival Tree
Building a Survival Tree
Constructing a...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Methods for Analyzing Epidemiological Data

