Related Experiment Videos
Judging the quality of gene expression-based clustering methods using gene annotation
Francis D Gibbons1, Frederick P Roth
1Department of Biological Chemistry and Molecular Pharmacology, Harvard Medical School, Boston, Massachusetts 02115, USA.
Genome Research
|October 9, 2002
Summary
This study evaluated gene clustering algorithms using mutual information. Lower cluster numbers generally yield better biological function enrichment, with specific distance metrics optimal at different cluster counts.
Area of Science:
- Bioinformatics
- Computational Biology
- Systems Biology
Background:
- Gene expression data analysis is crucial for understanding biological systems.
- Clustering algorithms are widely used to group genes with similar expression patterns.
- Evaluating the performance of these algorithms is essential for reliable biological interpretation.
Purpose of the Study:
- To compare the performance of various gene clustering algorithms.
- To assess the impact of different distance metrics on clustering results.
- To determine optimal cluster numbers for biological function enrichment.
Main Methods:
- Utilized a figure of merit based on mutual information between cluster membership and gene attributes.
- Compared Euclidean and Pearson distances for ratio-based and non-ratio-based measurements.
- Evaluated algorithms including self-organized maps and hierarchical clustering (single- and average-linkage).
- Analyzed multiple publicly available gene expression datasets.
Main Results:
- Biological function enrichment is generally highest at lower cluster numbers.
- Euclidean distance is optimal for ratio-based measurements, and Pearson distance for non-ratio-based measurements at optimal cluster numbers.
- Self-organized map (SOM) approach performs best for both measurement types at higher cluster numbers.
- Hierarchical clustering methods (single- and average-linkage) yielded worse-than-random results.
Conclusions:
- The choice of clustering algorithm and distance metric significantly impacts the biological interpretability of gene expression data.
- Lower cluster numbers often lead to more biologically meaningful gene groupings.
- Self-organized maps offer a robust approach for gene clustering, particularly at higher cluster counts.