Related Experiment Videos
Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional
Johannes A Vey1, Georg Heinze2, Meinhard Kieser3
1Institute of Medical Biometry, University of Heidelberg, Im Neuenheimer Feld 130.3, Heidelberg, 69120, Germany. vey@imbi.uni-heidelberg.de.
Purpose:
A wide range of methods exist for developing a clinical prediction model (CPM) and for performing variable selection. Our purpose was to develop a fair simulation study design and to investigate the properties, strengths, and weaknesses of different methods to predict a continuous outcome in low-dimensional data situations.
Methods:
In this simulation study, we conducted a neutral comparison of traditional (linear regression with stepwise selection) and machine learning (regularized regression with elastic net, gradient boosting, random forest) variable selection strategies to derive a CPM. The generated datasets included a total of 15 variables, with 8 of those being predictor variables. Four data- and outcome-generating mechanisms with increasing complexity produced data structures typical for biomedicine covering linear associations and gradually introducing non-linear and non-additive elements into the data structure.
Results:
All methods generally performed better with increasing sample size and less noise in the data. Gradient boosting with regression models and with trees as base learners, and the elastic net regularized regression included nearly all variables (i.e., both the predictor and non-predictor variables), especially with increasing sample size. The linear regression model with stepwise selection (LMSS) showed the best trade-off between correctly including the predictors and excluding the non-predictor variables in most of the scenarios, even when the functional form of continuous predictors deviated from linearity. In more complex data, variable selection using the Boruta or Hapfelmeier approach for random forest performed similar to LMSS.
Conclusion:
The sample size must be sufficiently large to enable the methods to reliably identify the predictor variables and to ensure that the developed CPMs are accurate and well-calibrated. LMSS revealed good properties and the random forest with the Boruta or Hapfelmeier approach are suitable alternatives if complex associations between predictors and outcomes are assumed.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.