Related Experiment Videos
Clustering protein sequences--structure prediction by transitive homology
E Bolten1, A Schliep, S Schneckener
1Institut für Biochemie, Universität zu Köln, Weyertal 80, D-50937 Köln, Germany. e.bolten@science-factory.com
Bioinformatics (Oxford, England)
|October 24, 2001
Summary
This study introduces a novel graph-based clustering method to identify remote protein homologues. The approach leverages transitivity and improves remote homology detection by 24% compared to traditional pair-wise comparisons.
Area of Science:
- Bioinformatics
- Computational Biology
- Structural Bioinformatics
Background:
- Protein sequence identity above a threshold implies structural similarity and common ancestry.
- Remote homologue detection requires criteria beyond simple sequence identity.
- Transitivity, inferring similarity via intermediate proteins, is explored for homology detection.
Purpose of the Study:
- To develop and evaluate a graph-based clustering approach for identifying remote protein homologues.
- To investigate the role and limitations of transitivity in protein homology detection.
- To improve the accuracy and efficiency of detecting structural similarity from sequence data.
Main Methods:
- Utilized Smith-Waterman algorithm for all-pair sequence similarity in SwissProt.
- Constructed a directed graph where nodes are protein sequences and edges represent scaled similarity.
- Developed a novel graph-based clustering algorithm incorporating transitivity and handling multi-domain proteins.
Main Results:
- The graph-based clustering method effectively uses transitivity for remote homologue detection.
- Length-dependent scaling of alignment scores prevents clustering errors in multi-domain proteins.
- Achieved a 24% improvement in detecting remote homologues compared to pair-wise comparisons using SCOP as a benchmark.
Conclusions:
- The developed graph-based clustering approach enhances remote homology detection.
- Transitivity, when appropriately constrained, is a valuable principle for inferring protein relationships.
- The method offers an efficient way to analyze large protein sequence datasets for structural similarity.