Related Experiment Videos
Tumor classification and marker gene prediction by feature selection and fuzzy c-means clustering using microarray
Junbai Wang1, Trond Hellem Bø, Inge Jonassen
1Department of Tumor Biology, The Norwegian Radium Hospital, N0310 Oslo, Norway. junbaiw@radium.uio.no
BMC Bioinformatics
|December 4, 2003
Summary
This study introduces novel DNA microarray models for accurate tumor classification and marker gene prediction, significantly reducing error rates and offering potential for improved medical diagnostics.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Development of novel computational models for tumor classification and gene prediction using DNA microarrays.
- Utilizes Self-Organizing Maps (SOMs) for summarizing gene expression profiles.
- Employs Fuzzy C-means clustering for tumor sample classification.
Purpose of the Study:
- To develop and validate novel models for accurate tumor classification and prediction of marker genes from DNA microarray data.
- To assess the performance of the proposed models against existing methods.
- To highlight the importance of feature selection in microarray data analysis.
Main Methods:
- Gene expression data summarization using optimally selected Self-Organizing Maps (SOMs).
- Tumor sample classification via Fuzzy C-means clustering.
- Marker gene prediction using manual (SOM component plane visualization) or automatic (Fisher's linear discriminant) feature selection.
Main Results:
- Models tested on four diverse datasets: Leukemia, Colon cancer, Brain tumors, and NCI cancer cell lines.
- Achieved significantly reduced error rates in class prediction compared to other approaches.
- Demonstrated the critical role of feature selection in enhancing microarray data analysis.
Conclusions:
- The developed models effectively identify marker genes with high predictive potential, outperforming existing methods.
- Models show promise for medical diagnostics and offer insights into cancer classification.
- Identified inherent limitations in tumor classification from microarray data related to data class size and internal class structure.