Related Experiment Video
Updated: Jul 14, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Selecting dissimilar genes for multi-class classification, an application in cancer subtyping.
Zhipeng Cai1, Randy Goebel, Mohammad R Salavatipour
1Department of Computing Science, University of Alberta, Edmonton, Alberta, Canada. zhipeng@cs.ualberta.ca <zhipeng@cs.ualberta.ca>
BMC Bioinformatics
|June 19, 2007
Summary
This study introduces a new method for cancer subtyping using gene expression data. It improves classification accuracy by selecting less correlated genes, outperforming previous methods in identifying cancer subtypes.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Gene expression microarrays are vital for disease profiling and treatment strategies.
- Biomarker identification is crucial for constructing accurate diagnostic and subtype classification models.
- While binary classification in microarray data is established, multi-class classification, like cancer subtyping, remains a significant challenge.
Purpose of the Study:
- To address the challenges in multi-class classification of microarray data, specifically for cancer subtyping.
- To develop a novel approach for gene representation and selection to enhance classifier performance.
- To improve the accuracy of cancer subtyping using gene expression data.
Main Methods:
- Introduction of a novel class discrimination strength vector for individual gene representation.
- Development of a new metric to quantify class discrimination strength differences between genes.
- Application of this metric in gene clustering for selecting informative, less correlated genes for classifier construction.
Main Results:
- Tested on four real-world cancer microarray datasets, the proposed method significantly improved classification accuracy.
- The constructed classifiers outperformed previously reported best results on these datasets.
- Selected genes were less correlated and contributed statistically significantly to more accurate cancer subtyping.
Conclusions:
- The novel class discrimination strength vector offers a superior gene representation compared to standard gene expression vectors.
- This method effectively eliminates redundant, highly correlated genes, leading to more robust classifier construction.
- The approach demonstrated enhanced accuracy in cancer subtyping, highlighting its potential in clinical applications.
