Related Experiment Videos
Graph-based clustering for finding distant relationships in a large set of protein sequences
Hideya Kawaji1, Yoichi Takenaka, Hideo Matsuda
1Department of Bioinformatic Engineering, Graduate School of Information Science and Technology, Osaka University, 1-3 Machikaneyama, Toyonaka, Osaka 560-8531, Japan.
Bioinformatics (Oxford, England)
|January 22, 2004
Summary
A new graph-based algorithm efficiently clusters distantly-related proteins by partitioning sequence similarity networks. This method achieves high recall, identifying novel protein relationships beyond traditional clustering techniques.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Protein sequence clustering is crucial for functional characterization.
- Clustering distantly-related proteins with only regional sequence similarity remains challenging.
- Developing advanced algorithms for distantly-related protein clustering is necessary.
Purpose of the Study:
- To develop a time and space efficient algorithm for clustering distantly-related proteins.
- To improve the accuracy and scope of protein functional characterization through advanced clustering.
Main Methods:
- Developed a graph-based clustering algorithm.
- Represented proteins as vertices and sequence similarities as weighted edges.
- Employed a novel score combining normalized cut and locally minimal cut capacities for graph partitioning.
- Applied the algorithm to 40,703 human proteins from SWISS-PROT and TrEMBL.
Main Results:
- Achieved 76% recall for 20,529 proteins compared to InterPro classifications.
- Successfully identified protein relationships missed by other clustering methods.
- Demonstrated efficiency in both time and space.
Conclusions:
- The developed algorithm effectively clusters distantly-related proteins.
- This method enhances the functional characterization of large protein datasets.
- The algorithm offers a valuable tool for bioinformatics research.