Related Experiment Video
Updated: Mar 31, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A tutorial on variable selection for clinical prediction models: feature selection methods in data mining could
Farideh Bagherzadeh-Khiabani1, Azra Ramezankhani1, Fereidoun Azizi2
1Prevention of Metabolic Disorders Research Center, Research Institute for Endocrine Sciences, Shahid Beheshti University of Medical Sciences, Velenjak, 1985717413 Tehran, Iran.
Data mining variable selection methods significantly improve clinical prediction models for diabetes. Wrapper methods performed best, enhancing model accuracy and efficiency in epidemiological studies.
Area of Science:
- Epidemiology
- Clinical Prediction
- Data Mining
Background:
- Predictor selection is crucial for accurate clinical prediction models.
- Traditional epidemiological methods may not be optimal for complex datasets.
Purpose of the Study:
- To evaluate data mining variable selection techniques in an epidemiological context.
- To compare the performance of filter and wrapper methods against traditional approaches.
Main Methods:
- Applied P-value, filter (symmetrical uncertainty), and wrapper methods to select diabetes predictors.
- Utilized logistic regression on a training set and evaluated on a test set.
- Performance assessed using Akaike Information Criterion (AIC) and Area Under the Curve (AUC).
Main Results:
- Wrapper-based models yielded the best performance.
- Symmetrical uncertainty filter method also showed strong results for AUC and AIC.
- The full model with all variables performed worst.
Conclusions:
- Data mining variable selection methods enhance clinical prediction model performance.
- Wrapper methods are particularly effective for predictor selection in epidemiological studies.
- An R program facilitates the application and visualization of these methods.
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Statistical Software for Data Analysis and Clinical Trials
Survival Tree
Building a Survival Tree
Constructing a...
Kaplan-Meier Approach
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Clinical Trials
There are four phases in a clinical trial. A phase one...

