Related Experiment Videos
Selection bias in gene extraction on the basis of microarray gene-expression data.
Christophe Ambroise1, Geoffrey J McLachlan
1Laboratoire Heudiasyc, Unité Mixte de Recherche/Centre National de la Recherche Scientifique 6599, 60200 Compiègne, France.
Summary
Accurate cancer prediction rules require accounting for selection bias. Correcting for this bias reveals that prediction error rates are not negligible when using only a few genes.
Area of Science:
- Bioinformatics
- Cancer Genomics
- Biostatistics
Background:
- Developing accurate cancer prediction rules from gene expression data is challenging due to high dimensionality and limited sample sizes.
- Previous studies suggested low prediction error rates using small gene subsets, but often overlooked selection bias.
Purpose of the Study:
- To assess and correct for selection bias in constructing gene expression-based cancer prediction rules.
- To evaluate the impact of bias correction on prediction error rates.
Main Methods:
- Utilized cross-validation and bootstrap methods external to the gene selection process.
- Recommended 10-fold cross-validation and the .632+ bootstrap error estimate.
- Applied methods to two published cancer gene expression datasets.
Main Results:
- Selection bias inflates the accuracy of prediction rules, leading to underestimated error rates.
- After correcting for selection bias, cross-validated error rates were non-zero even for small gene subsets.
- The .632+ bootstrap method effectively handles overfitted prediction rules.
Conclusions:
- Failure to account for selection bias can lead to overly optimistic assessments of prediction rule performance in cancer genomics.
- Robust validation strategies, like external cross-validation or appropriate bootstrap methods, are crucial for reliable cancer prediction models.
- The choice of validation technique significantly impacts the perceived accuracy of gene-based diagnostic tools.