Related Experiment Videos
Correcting the loss of cell-cycle synchrony in clustering analysis of microarray data using weights
1Department of Epidemiology and Public Health, Yale University School of Medicine, New Haven, CT 06520-8034, USA.
Bioinformatics (Oxford, England)
|May 29, 2004
Summary
Weighted k-means clustering improves analysis of cell-cycle data by accounting for temporal patterns. This method enhances the agreement between gene expression clusters and known protein complexes, aiding biological discovery.
Area of Science:
- Computational Biology
- Systems Biology
- Bioinformatics
Background:
- Standard clustering methods like k-means struggle with cell-cycle data due to loss of synchrony.
- These methods group genes based solely on expression levels, ignoring crucial temporal dynamics.
Purpose of the Study:
- To enhance the performance of k-means clustering for time-series gene expression data.
- To improve the biological relevance of clustering by incorporating temporal information.
Main Methods:
- A 'weighted k-means' approach was developed, assigning decreasing weights to variables over time.
- The method was evaluated using a yeast cell-cycle dataset.
- Adjusted Rand index was used to compare clustering results with known protein complex structures.
Main Results:
- The proposed time-decreasing weight function, exp[-(1/2)(t(2)/C(2))], significantly improved clustering accuracy.
- Optimal performance was observed when the parameter C approximated the length of two cell cycles.
- Increased agreement was found between k-means clusters and biological protein complexes.
Conclusions:
- Weighted k-means effectively addresses the limitations of standard clustering for asynchronous cell-cycle data.
- Incorporating temporal weighting enhances the biological interpretability of gene expression clustering.
- This approach offers a more robust method for analyzing dynamic biological processes.