Related Experiment Video
Updated: Jul 3, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Robust and efficient identification of biomarkers by classifying features on graphs
TaeHyun Hwang1, Hugues Sicotte, Ze Tian
1Department of Computer Science and Engineering, University of Minnesota, Twin Cities, Bioinformatics Core, Mayo Clinic College of Medicine, Rochester, MN, USA.
Bioinformatics (Oxford, England)
|July 26, 2008
Summary
This study introduces a novel graph-based algorithm for biomarker discovery from gene expression and SNP data. The method enhances reproducibility and handles large datasets, identifying relevant disease markers.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Biomarker discovery from large-scale gene expression or single nucleotide polymorphism (SNP) data faces computational challenges due to feature dependence.
- Ignoring feature dependence leads to non-reproducible biomarkers across datasets.
- High-throughput data analysis requires scalable and robust methods for accurate biomarker identification.
Purpose of the Study:
- To introduce a novel graph-based semi-supervised feature classification algorithm for identifying discriminative disease markers.
- To capture feature dependence using network propagation on bipartite graphs.
- To achieve reproducible biomarker identification and handle large-scale datasets.
Main Methods:
- Developed a graph-based semi-supervised feature classification algorithm utilizing bipartite graphs.
- Employed network propagation to capture dependencies among samples and features (clinical and genetic variables).
- Explored bi-cluster structures within the graph for enhanced feature classification.
Main Results:
- Applied the network propagation algorithm to three large-scale breast cancer datasets.
- Achieved competitive classification performance compared to Support Vector Machines (SVMs) and other baseline methods.
- Identified clinically and biologically relevant markers, including highly reproducible marker genes and enriched functions.
Conclusions:
- The developed algorithm effectively identifies reproducible biomarkers from high-throughput data.
- The method demonstrates scalability for handling hundreds of thousands of features.
- The algorithm can simultaneously classify features and test samples for disease prognosis and diagnosis.
