Related Experiment Video
Updated: Jun 2, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Comparison of Random Forest and Stepwise Regression for Variable Selection Using Low Prevalence Predictors: A case
Patricia Gilholm1, Paula Lister2,3,4, Adam Irwin5,6
1Children's Intensive Care Research Program, Child Health Research Centre, The University of Queensland, Brisbane, QLD, Australia. p.gilholm@uq.edu.au.
Random Forest and stepwise regression both effectively select variables, even low prevalence ones, in clinical prediction models. Both methods showed comparable predictive performance in a paediatric sepsis screening tool study.
Area of Science:
- Clinical Informatics
- Biostatistics
- Machine Learning in Healthcare
Background:
- Variable selection is crucial for identifying predictive factors in clinical data.
- Low prevalence predictors (LPPs) are common but understudied in variable selection.
- This study addresses LPPs in a paediatric sepsis screening tool.
Purpose of the Study:
- Compare Random Forest (RF) and stepwise regression (SWR) for variable selection.
- Evaluate the impact of LPPs on model performance.
- Assess predictor selection with varying prevalence thresholds.
Main Methods:
- Compared RF against forward and backward SWR for variable selection.
- Utilized a paediatric sepsis screening dataset with numerous LPPs.
- Assessed model performance via Area Under the Curve (AUC) and retained variables.
- Conducted a simulation study on predictor prevalence impact.
Main Results:
- RF retained 22 predictors (14 LPPs), SWR retained 17 (10 LPPs).
- Both models achieved similar predictive performance (RF AUC: 0.79, SWR AUC: 0.80).
- Simulation showed differing variable importance and selection with increasing prevalence thresholds for both methods.
Conclusions:
- RF selected more LPPs than SWR, but predictive performance was comparable.
- Both RF and SWR are suitable for variable selection with LPPs when applied correctly.
- Model performance is robust even with a substantial number of LPPs in small candidate predictor sets.
More Related Videos
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Comparing the Survival Analysis of Two or More Groups
Biostatistics: Overview
Discrete variables are...
Randomized Experiments
Simple randomization
Simple...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Censoring Survival Data

