Related Experiment Video
Updated: Jun 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Use of wrapper algorithms coupled with a random forests classifier for variable selection in large-scale genomic
Andrei S Rodin1, Anatoliy Litvinenko, Kathy Klos
1Human Genetics Center, School of Public Health, University of Texas Health Science Center, Houston, Texas, USA. asrodin@hotmail.com
Variable selection in large genetic studies is crucial. A new wrapper strategy using Random Forests efficiently identifies key single-nucleotide polymorphisms (SNPs) for coronary heart disease and LDL cholesterol prediction.
Area of Science:
- Genetics
- Bioinformatics
- Computational Biology
Background:
- Large-scale genetic association studies produce high-dimensional data, necessitating variable selection for computational efficiency and to prevent overfitting.
- Traditional data analysis methods struggle with the scale of modern genomic datasets.
Purpose of the Study:
- To introduce and evaluate a novel "wrapper" strategy (SIZEFIT) for efficient variable selection in genetic association studies.
- To identify a minimal set of single-nucleotide polymorphisms (SNPs) that retain predictive accuracy for complex traits like coronary heart disease (CHD) and low-density lipoprotein (LDL) cholesterol levels.
- To compare the performance of the proposed method with a statistical approach (SUMSTAT) and refine SNP subsets using a novel method (FIXFIT).
Main Methods:
- A Random Forests classifier was employed within a "wrapper" variable selection strategy (SIZEFIT).
- The methods were applied to a large case-cohort dataset (Atherosclerosis Risk in Communities) with 2,425 individuals and 4,869 SNPs.
- Novel refinement method (FIXFIT) using Extremal Optimization was developed for signal-containing SNPs.
Main Results:
- Most SNPs could be removed without compromising predictive accuracy, with predictive signals concentrated in a small subset of SNPs (often <100).
- Compact optimal predictive SNP subsets were constructed for CHD (<150 SNPs) and LDL (<300 SNPs).
- Significant overlap was observed between SNP rankings generated by different methods, suggesting robustness.
Conclusions:
- The SIZEFIT wrapper strategy is effective for reducing dimensionality in genome-wide association datasets while preserving predictive power.
- A small number of SNPs contain substantial predictive information for CHD and LDL cholesterol.
- The findings provide practical guidelines for generating compact, predictive SNP subsets from large genetic datasets.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Evolutionary Relationships through Genome Comparisons
Randomized Experiments
Simple randomization
Simple...
Multiple Allele Traits
Multiple Allele Traits
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...

