Related Experiment Video
Updated: May 24, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.6K
Automated sparse feature selection in high-dimensional proteomics data via 1-bit compressed sensing and K-Medoids
FuDong Wen1, Yue Su1, Dan Liu1
1Department of Biostatistics, Public Health College, Harbin Medical University, Harbin City, 150081, Heilongjiang Province, China.
BMC Bioinformatics
|July 2, 2025
Summary
Soft-Thresholded Compressed Sensing (ST-CS) automates biomarker discovery in high-dimensional proteomics. This novel method improves feature selection accuracy and classification performance, offering a more efficient approach for identifying key biomarkers.
Area of Science:
- Proteomics
- Biomarker Discovery
- Computational Biology
Background:
- High-dimensional proteomics data present challenges in biomarker discovery, including noise, redundancy, and multicollinearity.
- Existing feature selection methods lack stability, sparsity, and computational efficiency for complex proteomic datasets.
- Manual thresholding in conventional methods can lead to suboptimal biomarker identification.
Purpose of the Study:
- To introduce Soft-Thresholded Compressed Sensing (ST-CS), a hybrid framework for automated feature selection in high-dimensional proteomics.
- To address the limitations of current methods in stability, sparsity, and computational efficiency.
- To dynamically partition coefficient magnitudes for accurate discrimination between biomarkers and noise.
Main Methods:
- Integration of 1-bit compressed sensing with K-Medoids clustering.
- Automated feature selection through dynamic partitioning of coefficient magnitudes.
- Dynamic thresholding based on coefficient magnitude distribution.
Main Results:
- ST-CS demonstrated superior feature selection robustness with high sensitivity and specificity, reducing false discovery rates by 20-50% compared to Hard-Thresholded Compressed Sensing (HT-CS).
- Achieved superior F1 scores and Matthews Correlation Coefficients (MCC), outperforming HT-CS, LASSO, and SPLSDA in identifying true biomarkers.
- Outperformed all compared methods in classification accuracy (AUC) across varying noise levels while maintaining sparsity, with significant reductions in selected features on CPTAC datasets.
Conclusions:
- ST-CS rigorously automates feature selection in high-dimensional proteomics.
- The framework effectively balances classification efficacy, interpretability, and scalability for translational proteomics.
- ST-CS offers a powerful and efficient tool for precise biomarker discovery and predictive modeling.

