Related Experiment Video
Updated: Feb 11, 2026

Selective Capture of 5-hydroxymethylcytosine from Genomic DNA
Published on: October 5, 2012
Benchmarking Sparse Variable Selection Methods for Genomic Data Analyses
Hema Sri Sai Kollipara1, Tapabrata Maiti1, Sanjukta Chakraborty2
1Department of Statistics & Probability, Michigan State University, East Lansing, Michigan, USA.
This study compares Bayesian variable selection methods for genomic analysis. No single method excels across all scenarios, but LASSO, spike-and-slab (SN), and RFSFS show strong performance, especially with correlated features.
Area of Science:
- Genomics
- Statistical Genetics
- Computational Biology
Background:
- Genomic studies involve numerous features, necessitating accurate variable selection.
- Bayesian inference has advanced for variable selection, but practical implementation details and performance comparisons are lacking.
Purpose of the Study:
- To conduct a comparative analysis of Bayesian variable selection approaches for genomic data.
- To evaluate the performance of shrinkage, global-local, mixture priors, SUSIE, and a proposed RFSFS method.
Main Methods:
- Comparative analysis of Bayesian variable selection methods.
- Evaluation using metrics like False Discovery Rate (FDR), False Negative Rate (FNR), F-score, and mean squared prediction error.
- Simulation studies under various scenarios, including uncorrelated and correlated features.
Main Results:
- No single method uniformly outperforms others across all scenarios and metrics.
- LASSO, spike-and-slab prior with normal slab (SN), and RFSFS are competitive for FDR and F-score with uncorrelated features.
- SN, SuSIE, and RFSFS are competitive for FDR with correlated features; LASSO excels in F-score over SuSIE.
Conclusions:
- Method performance varies depending on feature correlation and evaluation metrics.
- The proposed RFSFS method demonstrates competitive performance alongside established techniques like LASSO and SN.
- Findings offer methodological direction for variable selection in genomic analyses, including The Cancer Genome Atlas (TCGA) data.
Related Concept Videos
Selected Data About Geographic Locations
Genomics
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Random Variables
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Statistical Methods for Analyzing Epidemiological Data
Graphs of Equations in Two Variables

