Related Experiment Videos
Is cross-validation better than resubstitution for ranking genes?
Ulisses Braga-Neto1, Ronaldo Hashimoto, Edward R Dougherty
1Section of Clinical Cancer Genetics, University of Texas M. D. Anderson Cancer Center, Houston, TX, USA.
Bioinformatics (Oxford, England)
|January 22, 2004
Summary
Resubstitution and cross-validation methods for ranking gene feature sets were compared. Resubstitution offers competitive ranking accuracy with significant computational savings compared to cross-validation.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning
Background:
- Ranking gene feature sets is crucial for phenotype classification and genetic regulatory network prediction.
- Resubstitution and cross-validation are common methods for estimating classifier error, but differ in bias and variability.
- Accurate ranking of gene sets is often more important than precise error estimation.
Purpose of the Study:
- To compare the ranking performance of resubstitution and cross-validation for gene feature set selection.
- To evaluate these methods in both classification and prediction tasks.
- To assess computational efficiency alongside ranking accuracy.
Main Methods:
- A model-based approach was employed to compare ranking performances.
- Methods included classification using Gaussian models, linear discriminant analysis, and 3-nearest-neighbor rules.
- Prediction was analyzed within probabilistic Boolean networks (PBNs).
Main Results:
- Resubstitution demonstrated competitive ranking accuracy compared to cross-validation across all tested scenarios.
- The model-based approach allowed direct comparison of ranking based on estimated versus true error.
- Resubstitution provided substantial computational time savings.
Conclusions:
- Resubstitution is a viable and computationally efficient alternative to cross-validation for ranking gene feature sets.
- The findings support the use of resubstitution when ranking accuracy is the primary concern.
- This study offers practical insights for gene set selection in bioinformatics and systems biology.