Related Experiment Video
Updated: Jun 13, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
The use of variable selection in clinical prediction modeling for binary outcomes: a systematic review
Xinrui Su1, Gareth Ambler1, Nathan Green1
1Department of Statistical Science, UCL, London, UK.
Objectives:
Clinical prediction models are valuable tools that can support medical decision-making. Concise and interpretable models are more likely to be adopted in clinical practice; therefore, appropriate selection of predictor variables is often considered essential in model development. Typically, researchers specify a list of candidate predictors based on literature reviews and expert knowledge. Data-driven variable selection methods are then often used to further reduce the number of variables in the final model. However, many commonly used approaches, such as univariable selection, have been generally discouraged in prediction modeling. This systematic review aims to examine current practice regarding the use of data-driven variable selection when developing clinical prediction models for binary outcomes using logistic regression.
Study Design And Setting:
We focused on published articles in PubMed between October 1 and October 21, 2024, that developed prediction models for binary health outcomes using logistic regression. We extracted information on the study characteristics and, if applicable, the methodology used for variable selection.
Results:
In total, 141 studies were included in the review. We found that nearly all studies (140/141) used data-driven variable selection. Univariable selection was by far the most used method; it was used in 78% (110/141) of studies. Other frequently used methods included backward elimination (BE) (60/141, 43%), "bulk removal" of variables (BR) from a single multivariable model (58/141, 41%), and Least Absolute Shrinkage and Selection Operator (LASSO) (35/141, 25%). Many studies applied a sequential application of variable selection methods; the most common two-step combinations were univariable selection followed by BE (45/139, 32%) and univariable selection followed by BR (43/139, 31%). In addition, many studies lacked sufficient detail in their reporting. Common problems included incomplete reporting of candidate predictors, and unclear specification of variable selection methods.
Conclusion:
Although data-driven variable selection is generally discouraged in clinical prediction modeling, nearly all studies in our review employed at least one such method, with many studies using two methods. Some of the most frequently criticized methods such as univariable selection and BE were commonly used. Modern penalized methods such as LASSO, which directly aim to optimize out-of-sample predictive performance while also removing redundant variables, were used less frequently.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Biostatistics: Overview
Discrete variables are...
Receiver Operating Characteristic Plot