Fuzzy c-means clustering with prior biological knowledge.
Luis Tari1, Chitta Baral, Seungchan Kim
1School of Computing and Informatics, Department of Computer Science and Engineering, Ira A. Fulton School of Engineering, Arizona State University, P.O. Box 878809, Tempe, AZ 85287-8809, USA. luis.tari@asu.edu
We developed GO Fuzzy c-means, a novel semi-supervised clustering method integrating biological knowledge and gene expression data. This approach enhances the biological relevance of gene clusters and improves gene function prediction.
Area of Science:
- Bioinformatics
- Computational Biology
- Systems Biology
Background:
- Gene expression data analysis often relies on clustering to identify groups of functionally related genes.
- Traditional clustering methods may not fully capture the complex, multi-functional nature of genes.
- Integrating prior biological knowledge can improve the accuracy and interpretability of clustering results.
Purpose of the Study:
- To introduce a novel semi-supervised clustering algorithm, GO Fuzzy c-means, that leverages biological knowledge.
- To enable probabilistic assignment of genes to multiple clusters, reflecting their potential multi-functionality.
- To enhance the biological meaningfulness of gene clusters and improve gene function prediction.
Main Methods:
- The GO Fuzzy c-means method is based on the fuzzy c-means clustering algorithm.
- It utilizes Gene Ontology (GO) annotations as prior biological knowledge to guide gene grouping.
- The algorithm allows genes to belong to multiple clusters, offering a more nuanced representation.
Main Results:
- GO Fuzzy c-means demonstrated superior performance in generating biologically meaningful clusters compared to state-of-the-art methods.
- The method effectively produced relevant clusters even with limited Gene Ontology annotations.
- Prior knowledge integration in GO Fuzzy c-means proved effective for predicting gene functions.
Conclusions:
- GO Fuzzy c-means offers a robust approach for semi-supervised gene clustering by integrating biological knowledge.
- The method provides a more accurate representation of gene behavior through multi-cluster assignments.
- This approach significantly improves the biological interpretability and functional prediction capabilities of gene expression data analysis.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Vesicular Tubular Clusters
With the help of motor proteins such...
T Cell Activation and Clonal Selection
Naive T cells that have not yet encountered an antigen express two primary CD...
Applications of Molecular Taxonomy
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order to...
Phylogenetic Trees
