Related Experiment Videos
A prediction-based resampling method for estimating the number of clusters in a dataset
Sandrine Dudoit1, Jane Fridlyand
1Division of Biostatistics, School of Public Health, University of California Berkeley, 140 Earl Warren Hall, Berkeley, CA 94720-7360, USA. sandrine@stat.berkeley.edu
Genome Biology
|August 20, 2002
Summary
A new method called Clest accurately estimates the number of tumor classes in gene-expression data. This prediction-based resampling technique is more robust than existing methods for cancer microarray studies.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Microarray technology is vital in biological and medical research, particularly for tumor classification.
- Identifying new tumor classes from gene-expression profiles presents a significant statistical challenge.
- Estimating the number of clusters and assigning samples are key aspects of this problem.
Purpose of the Study:
- To develop and evaluate a novel method for estimating the number of clusters in datasets.
- To address the challenge of determining the optimal number of tumor classes in gene-expression data.
- To improve the accuracy and robustness of cluster number estimation in cancer research.
Main Methods:
- Development of a prediction-based resampling method named Clest.
- Comparison of Clest against six existing methods using simulated and real-world gene-expression data.
- Evaluation of method performance on four published cancer microarray studies.
Main Results:
- Clest demonstrated superior accuracy and robustness compared to existing methods.
- The new method effectively estimates the number of clusters in gene-expression datasets.
- Performance was validated across diverse cancer microarray datasets.
Conclusions:
- Combining prediction accuracy with resampling yields reliable estimates for the number of clusters.
- Clest offers an accurate and robust solution for a critical problem in cancer genomics.
- The findings support the use of Clest for analyzing gene-expression data in tumor classification.