Related Experiment Video
Updated: Jul 10, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Prediction potential of candidate biomarker sets identified and validated on gene expression data from multiple
Michael Gormley1, William Dampier, Adam Ertel
1School of Biomedical Engineering, Drexel University, Philadelphia, PA, USA. mpg33@drexel.edu
BMC Bioinformatics
|October 30, 2007
Summary
Gene expression profiles accurately predict molecular phenotypes but show little agreement across different datasets. Predicting relapse directly from microarray data using machine learning is challenging.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Gene expression profiles from independent studies of the same condition often lack common genes.
- Microarray datasets from breast, lymphoma, and renal cancer samples were analyzed.
- Machine learning and ROC curves were employed to assess profile prediction error.
Purpose of the Study:
- To investigate the consistency of gene expression profiles in predicting molecular phenotypes and clinical outcomes.
- To compare the efficacy of supervised feature selection against random and a priori gene selection.
- To evaluate the generalizability of expression profiles across independent datasets and microarray platforms.
Main Methods:
- Iterative machine learning algorithm to create expression profile populations.
- Receiver Operating Characteristic (ROC) curves for prediction error assessment.
- Comparison of profiles correlated with molecular phenotype versus relapse-free status.
- Supervised univariate feature selection compared to random and a priori gene selection.
- Cross-dataset validation on independent samples from same or different microarray platforms.
Main Results:
- Highly discriminative expression profiles were generated for molecular phenotypes (ER, BCL-6).
- Prognostic prediction using relapse-free status yielded poorly discriminative decision rules.
- Supervised feature selection outperformed random and a priori selection, with diminishing differences as feature numbers increased.
- Results were consistent when applied across datasets on the same microarray platform.
Conclusions:
- Numerous gene sets can accurately predict molecular phenotypes.
- Expect limited agreement between expression profiles derived from different training datasets.
- Direct prediction of relapse from microarray data using supervised machine learning is difficult.
- Findings are pertinent to identifying candidate biomarker panels through molecular profiling.
