Related Experiment Videos
Many accurate small-discriminatory feature subsets exist in microarray transcript data: biomarker discovery.
1Life Sciences Division, Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA. lesliegrate@comcast.net
BMC Bioinformatics
|April 14, 2005
Summary
Researchers found small gene sets (3 or less) that accurately classify microarray data, potentially serving as biomarkers for diagnostic tests. This exhaustive search method reveals previously unreported predictive gene combinations in cancer and leukemia datasets.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Molecular profiling generates high-dimensional gene expression data from biological samples.
- Classifying samples into two groups (e.g., normal vs. tumor) using this data is challenging.
- Existing methods identify small feature sets but lack guarantees on minimality.
Purpose of the Study:
- To investigate the existence of minimal, highly accurate gene classifiers in published microarray datasets.
- To explore the potential of small gene sets as biomarkers for disease diagnosis.
Main Methods:
- Employed a brute-force exhaustive search across all gene combinations (single genes, pairs, triples).
- Utilized a linear-hyperplane classification method to identify error-free classifiers.
- Applied the method to 10 diverse, published microarray datasets.
Main Results:
- All 10 datasets contained predictive small gene sets.
- Four datasets yielded thousands of gene pairs, while six had single genes for perfect discrimination.
- Discovered novel, accurate gene sets (≤3 genes) not reported in original publications.
Conclusions:
- Small gene sets capable of accurate classification are common in microarray data.
- These findings suggest potential for novel biomarkers and simple diagnostic tests.
- Identified specific gene sets for hepatocellular carcinoma (HCC) and leukemia subtypes, highlighting PLAC8 and GPC3.