Related Experiment Video
Updated: May 24, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Identifying a small set of marker genes using minimum expected cost of misclassification
Samuel H Huang1, Dengyao Mo, Jarek Meller
1School of Dynamic Systems, University of Cincinnati, 2600 Clifton Ave., Cincinnati, OH 45221, USA. sam.huang@uc.edu
Artificial Intelligence in Medicine
|March 6, 2012
Summary
A new feature selection method identifies minimal marker genes for predicting cancer and autoimmune disease phenotypes. This approach uses minimum expected cost of misclassification (MEMC) for superior accuracy with fewer features.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Accurate identification of marker genes is crucial for disease subtyping and phenotype prediction.
- Conventional feature selection methods may yield large gene sets, limiting clinical utility.
- Model-independent approaches offer potential for robust feature selection.
Purpose of the Study:
- To present a model-independent feature selection approach for identifying minimal marker gene subsets.
- To apply this method to predict breast cancer molecular subtypes (p53 status) and human leukocyte antigen (HLA) gene alleles.
Main Methods:
- Utilized minimum expected cost of misclassification (MEMC) to evaluate feature subset discriminative power without model building.
- Combined MEMC with sequential forward search for efficient feature selection.
- Applied the method to breast cancer and HLA genetic variant datasets.
Main Results:
- Identified two marker genes for p53 status with a p-value of 7.53×10(-5), outperforming previous methods.
- Selected six single-nucleotide polymorphism (SNP) loci for HLA phenotype prediction with 92.8% accuracy, surpassing other techniques.
- Demonstrated superior performance compared to traditional filter methods.
Conclusions:
- The MEMC-based feature selection effectively identifies smaller, highly discriminative marker gene sets.
- This approach offers improved performance and efficiency for molecular subtyping and phenotype prediction.
- The method holds promise for applications in precision medicine and disease research.

