Related Experiment Video
Updated: Aug 6, 2026

Basics of Multivariate Analysis in Neuroimaging Data
Published on: July 24, 2010
A preregistered simulation study provided evidence on the appropriate use of data-driven variable selection in
Theresa Ullmann1, Georg Heinze1, Michael Kammer2
1Institute of Clinical Biometrics, Center for Medical Data Science, Medical University of Vienna, Vienna, Austria.
Objectives:
Multivariable regression modeling is central in epidemiological research. Common practice often uses data-driven variable selection to identify the variables most strongly associated with the outcome (the goal of descriptive modeling), but this may lead to omission of relevant covariates, inclusion of noise variables, biased coefficient estimates, and unstable models. Although methodological recommendations exist, guidance based on systematic and neutral comparison studies remains limited. Therefore, we aimed to provide evidence-based guidance on the appropriate use of data-driven variable selection in epidemiological regression modeling, with a focus on descriptive modeling.
Study Design And Setting:
We conducted a comprehensive simulation study comparing commonly used variable selection approaches, including backward elimination (BE) with various selection criteria and variants of the least absolute shrinkage and selection operator (Lasso). Methods were evaluated across realistic epidemiological scenarios varying in outcome type (binary or continuous), sample size, and signal-to-noise ratio. Performance was assessed using measures relevant to descriptive modeling: variable selection rates (TPR: true positive rates and FPR: false positive rates), model selection rates, and bias of regression coefficients.
Results:
No single selection method performs uniformly well across all settings. Acceptable performance requires large sample sizes and/or strong signal-to-noise ratios. Under such favorable conditions, BE with a mild selection criterion (α = 0.5) performs best in terms of high TPR, BE with stricter criteria (α = 0.05 or Bayesian information criterion) performs best in terms of low FPR, and BE with Akaike information criterion strikes a balance. Lasso methods do not achieve optimal performance when evaluated against descriptive modeling criteria.
Conclusion:
Variable selection methods may complement substantive knowledge-driven specification but only under favorable conditions and not as the default strategy in epidemiological analyses. The choice of method should be guided by the desired trade-off between high TPR and low FPR, for which we provide a practical decision framework.
Plain Language Summary:
Researchers often use statistical models to study how different variables, such as age, gender, or medical conditions, are related to health outcomes. To develop these models, decisions have to be made about which variables to include and which to exclude. In practice, this is commonly done applying a data-driven variable selection method on the observed data. However, this approach can produce a misleading final statistical model. In this study, we examined how well popular data-driven variable selection methods are able to distinguish variables that are truly associated with the outcome from irrelevant noise variables. We used computer simulations that mimic real-world epidemiological data to compare different approaches, including commonly used techniques such as BE and Lasso. We evaluated these methods under a range of conditions, including small versus large sample sizes and strong versus weak relationships between variables and outcomes. As expected, we found that no method works well in all situations. In particular, all variable selection methods tended to perform poorly when the dataset was small or when the relationship between the variables and the outcome was weak. In these cases, important variables may be missed, irrelevant ones may be included, estimated effects can be biased, and the selected model can vary considerably depending on the specific data used. When conditions are more favorable (eg, with large datasets and strong relationships), variable selection methods can perform reasonably well. Simpler approaches such as BE can strike a reasonable balance between keeping important variables and excluding irrelevant ones, depending on how strict or lenient the statistical criteria for removing variables are. More complex methods, such as Lasso, did not consistently perform better. Overall, our results suggest that data-driven variable selection should not be used automatically. Instead, researchers should first rely on existing scientific knowledge to decide which variables to include. If data-driven variable selection is used subsequently, it should only be done under appropriate conditions and with a clear understanding of its limitations. With a decision framework, we provide practical guidance to help researchers decide when and how to use these methods more appropriately.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Biostatistics: Overview
Discrete variables are...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Regression Toward the Mean