Related Experiment Video
Updated: Nov 20, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Ranking of a wide multidomain set of predictor variables of children obesity by machine learning variable importance
Helena Marcos-Pasero1, Gonzalo Colmenarejo2, Elena Aguilar-Aguilar1
1Nutrition and Clinical Trials Unit, GENYAL Platform IMDEA-Food Institute, CEI UAM+CSIC, 28049, Madrid, Spain.
Insights
Machine learning models identified key factors contributing to childhood obesity in Spanish children aged 6-9. This research aids in developing targeted prevention strategies for pediatric obesity and related cardio-metabolic diseases.
Area of Science:
- Pediatrics
- Public Health
- Computational Biology
Background:
- Childhood obesity is a growing concern, linked to future cardio-metabolic diseases.
- Obesity's causes are complex, involving genetics, lifestyle, and environment, especially during rapid childhood growth.
- Machine learning (ML) is suitable for analyzing complex, high-dimensional data in obesity research.
Purpose of the Study:
- To utilize ML to identify significant predictors of body mass index (BMI) in children.
- To analyze a comprehensive set of 190 multidomain variables in a Spanish child cohort.
- To inform better prevention and treatment strategies for childhood obesity.
Main Methods:
- Analysis of 221 children (aged 6-9 years) from Madrid, Spain.
- Application of Random Forest and Gradient Boosting Machine models.
- Estimation of predictor importance using permutation and multiple imputation techniques.
Main Results:
- Identification of key variables associated with childhood obesity through ML analysis.
- Robust estimation of predictor importance from 190 multidomain factors.
- Insights into the complex interplay of factors influencing pediatric BMI.
Conclusions:
- ML models can effectively predict BMI and identify critical risk factors for childhood obesity.
- Understanding these factors is crucial for developing targeted interventions.
- This study provides a foundation for evidence-based childhood obesity prevention programs.
Abstract:
The increased prevalence of childhood obesity is expected to translate in the near future into a concomitant soaring of multiple cardio-metabolic diseases. Obesity has a complex, multifactorial etiology, that includes multiple and multidomain potential risk factors: genetics, dietary and physical activity habits, socio-economic environment, lifestyle, etc. In addition, all these factors are expected to exert their influence through a specific and especially convoluted way during childhood, given the fast growth along this period. Machine Learning methods are the appropriate tools to model this complexity, given their ability to cope with high-dimensional, non-linear data. Here, we have analyzed by Machine Learning a sample of 221 children (6-9 years) from Madrid, Spain. Both Random Forest and Gradient Boosting Machine models have been derived to predict the body mass index from a wide set of 190 multidomain variables (including age, sex, genetic polymorphisms, lifestyle, socio-economic, diet, exercise, and gestation ones). A consensus relative importance of the predictors has been estimated through variable importance measures, implemented robustly through an iterative process that included permutation and multiple imputation. We expect this analysis will help to shed light on the most important variables associated to childhood obesity, in order to choose better treatments for its prevention.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Outliers and Influential Points
Multi-input and Multi-variable systems
In the absence of...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Spearman's Rank Correlation Test
Spearman's test calculates correlation by...