Related Experiment Video
Updated: Aug 6, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Enhancing feature selection for ordinal outcomes using resampling-based sparse linear discriminant analysis
Yin Liu1,2, Ryan Wang3, Dong Si4
1Department of Neurobiology and Anatomy, McGovern Medical School, University of Texas Health Science Center at Houston, Houston, TX 77030, United States.
This study introduces a novel ensemble Sparse Linear Discriminant Analysis (sLDA) method using resampling to enhance feature selection stability. The approach improves reproducibility and identifies robust, interpretable biomarker signatures for biomedical data analysis.
Area of Science:
- Biostatistics
- Bioinformatics
- Computational Biology
Background:
- High-dimensional biomedical data with ordinal outcomes present feature selection challenges due to predictor correlations and small sample sizes.
- Sparse Linear Discriminant Analysis (sLDA) is used for classification and feature selection but suffers from instability and lack of reproducibility, especially with collinearity.
- Existing sLDA methods often fail to identify robust and interpretable biological marker sets.
Purpose of the Study:
- To develop a resampling-based ensemble sLDA framework to improve the stability and reproducibility of feature selection in high-dimensional biomedical data.
- To enhance the identification of biologically coordinated and interpretable marker signatures.
- To improve predictive performance and biomarker discovery in precision medicine.
Main Methods:
- A resampling-based ensemble Sparse Linear Discriminant Analysis (sLDA) framework integrating bootstrapping and subsampling was developed.
- Feature selection was based on Variable Inclusion Probability (VIP) aggregated across multiple resampled datasets.
- The proposed method was evaluated using simulation studies and applied to kidney renal papillary cell carcinoma staging and glioma grading datasets.
Main Results:
- The ensemble sLDA framework demonstrated improved stability and reproducibility of selected feature sets compared to standard sLDA.
- Simulation studies showed more accurate and consistent recovery of true predictors with the ensemble method.
- Applications revealed enhanced predictive performance and the identification of biologically interpretable and reproducible feature sets.
Conclusions:
- The proposed resampling-based ensemble sLDA framework offers a more stable and reproducible approach to feature selection for high-dimensional biomedical data.
- This method enhances biomarker discovery by identifying robust and interpretable feature sets, contributing to precision medicine.
- The framework's ability to handle collinearity and improve predictive accuracy makes it a valuable tool for analyzing complex biological datasets.
Related Concept Videos
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Ranks
Friedman Two-way Analysis of Variance by Ranks
Quantifying and Rejecting Outliers: The Grubbs Test
Outliers and Influential Points
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
