Ranking of a wide multidomain set of predictor variables of children obesity by machine learning variable importance

Helena Marcos-Pasero1, Gonzalo Colmenarejo2, Elena Aguilar-Aguilar1

  • 1Nutrition and Clinical Trials Unit, GENYAL Platform IMDEA-Food Institute, CEI UAM+CSIC, 28049, Madrid, Spain.

Scientific Reports
|January 22, 2021
PubMed

Insights

Machine learning models identified key factors contributing to childhood obesity in Spanish children aged 6-9. This research aids in developing targeted prevention strategies for pediatric obesity and related cardio-metabolic diseases.

Area of Science:

  • Pediatrics
  • Public Health
  • Computational Biology

Background:

  • Childhood obesity is a growing concern, linked to future cardio-metabolic diseases.
  • Obesity's causes are complex, involving genetics, lifestyle, and environment, especially during rapid childhood growth.
  • Machine learning (ML) is suitable for analyzing complex, high-dimensional data in obesity research.

Purpose of the Study:

  • To utilize ML to identify significant predictors of body mass index (BMI) in children.
  • To analyze a comprehensive set of 190 multidomain variables in a Spanish child cohort.
  • To inform better prevention and treatment strategies for childhood obesity.

Main Methods:

  • Analysis of 221 children (aged 6-9 years) from Madrid, Spain.
  • Application of Random Forest and Gradient Boosting Machine models.
  • Estimation of predictor importance using permutation and multiple imputation techniques.

Main Results:

  • Identification of key variables associated with childhood obesity through ML analysis.
  • Robust estimation of predictor importance from 190 multidomain factors.
  • Insights into the complex interplay of factors influencing pediatric BMI.

Conclusions:

  • ML models can effectively predict BMI and identify critical risk factors for childhood obesity.
  • Understanding these factors is crucial for developing targeted interventions.
  • This study provides a foundation for evidence-based childhood obesity prevention programs.

Related Concept Videos

Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.4K
Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
5.2K
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
258
Regression Analysis01:11

Regression Analysis

Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.0K
Variability: Analysis01:11

Variability: Analysis

Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
284
Spearman's Rank Correlation Test01:20

Spearman's Rank Correlation Test

Spearman's rank correlation test, also known as Spearman's rho, is a nonparametric method for assessing the strength and direction of association between two variables. This test is particularly valuable when the data distribution is unknown or when the assumption of normality does not hold. Named after the English psychologist and statistician Dr. Charles Edward Spearman, it serves as the nonparametric counterpart to Pearson's correlation coefficient.
Spearman's test calculates correlation by...
1.2K