Related Experiment Video
Updated: Dec 13, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Discovering a sparse set of pairwise discriminating features in high-dimensional data.
Samuel Melton1, Sharad Ramanathan2,3,4
1Applied Mathematics Harvard University, Cambridge, MA 02138, USA.
Discovering key features is essential for analyzing high-dimensional biological data. Our method identifies informative features in single-cell RNA sequencing data, revealing hidden cell states and improving developmental biology insights.
Area of Science:
- Computational Biology
- Developmental Biology
- Genomics
Background:
- High-dimensional biological data, such as single-cell RNA sequencing (scRNA-seq), offer unprecedented insights into complex processes like cell differentiation.
- Extracting meaningful biological insights and mechanistic understanding from these datasets remains a significant challenge.
- Identifying key molecular regulators driving developmental transitions from scRNA-seq data is particularly difficult.
Purpose of the Study:
- To develop an unsupervised method for identifying informative features in high-dimensional biological data.
- To uncover a low-dimensional subspace where clusters representing distinct biological states become linearly separable.
- To improve the analysis of scRNA-seq data and facilitate the discovery of key regulators in developmental processes.
Main Methods:
- Feature identification based on cluster discrimination.
- Development of an unsupervised method to find a low-dimensional subspace for clustering.
- Ensemble averaging of discriminators trained on proposed cluster configurations.
- Application to single-cell RNA-seq data from mouse gastrulation.
Main Results:
- Identification of 27 key transcription factors from 409 tested in mouse gastrulation scRNA-seq data.
- 18 of the identified transcription factors are known regulators of cell states.
- Discovery of a low-dimensional subspace revealing clear signatures of known cell types previously unclassifiable.
Conclusions:
- Identifying informative features is crucial for effective statistical analysis and predictive modeling in high-dimensional biological data.
- The proposed unsupervised method can uncover hidden biological structures in complex datasets.
- This approach enhances the ability to classify cell types and infer regulatory mechanisms in developmental biology.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Outliers and Influential Points
Causes of Similarity-Dissimilarity Effect
Wilcoxon Signed-Ranks Test for Matched Pairs
Expected Frequencies in Goodness-of-Fit Tests
Friedman Two-way Analysis of Variance by Ranks

