Related Experiment Video
Updated: Aug 9, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Variable Selection for Sparse Data with Applications to Vaginal Microbiome and Gene Expression Data.
Niloufar Dousti Mousavi1, Jie Yang1, Hani Aldirawi2
1Department of Mathematics, Statistics, and Computer Science, University of Illinois at Chicago, Chicago, IL 60607, USA.
This study introduces statistical methods for analyzing sparse, high-dimensional data. The methods successfully identified key differences in vaginal microbiome data and selected informative genes for accurate classification.
Area of Science:
- Statistics
- Bioinformatics
- Genomics
Background:
- Sparse data, characterized by a high proportion of zeros, is prevalent across scientific disciplines.
- Analyzing and modeling sparse, high-dimensional data presents significant statistical and computational challenges.
- Existing methods may not adequately address the complexities of real-world sparse datasets.
Purpose of the Study:
- To develop and present robust statistical methods and tools for analyzing sparse, high-dimensional data.
- To illustrate the application of these methods using real-world scientific datasets.
- To provide a framework for identifying significant features and improving predictive modeling in sparse data contexts.
Main Methods:
- Utilized zero-inflated model selection and significance testing for comparative analysis.
- Applied techniques to identify significant time intervals in longitudinal microbiome data.
- Employed feature selection methods to identify a subset of informative genes from a larger dataset.
Main Results:
- Successfully identified significant differences in *Lactobacillus* species between pregnant and non-pregnant groups.
- Selected the top 50 genes from 2426 sparse gene expression data, achieving 100% prediction accuracy.
- Demonstrated that the first four principal components of selected genes explain 83% of the model variability.
Conclusions:
- The proposed statistical methods are effective for analyzing complex sparse, high-dimensional data.
- The techniques facilitate the identification of biologically relevant features and improve classification accuracy.
- This approach offers a valuable tool for microbiome and gene expression data analysis.
Related Concept Videos
Genetic Variation
Genes exist in different versions called alleles,...
Quantifying and Rejecting Outliers: The Grubbs Test
DNA Microarrays
Frequency-dependent Selection
Statistical Software for Data Analysis and Clinical Trials
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...

