Related Experiment Videos
An efficient algorithm for large-scale detection of protein families.
A J Enright1, S Van Dongen, C A Ouzounis
1Computational Genomics Group, The European Bioinformatics Institute, EMBL Cambridge Outstation, Cambridge CB10 1SD, UK. anton@ebi.ac.uk
Nucleic Acids Research
|March 28, 2002
Summary
We developed TRIBE-MCL, a novel method for rapid and accurate protein family clustering using the Markov cluster algorithm. This approach effectively identifies protein families in large genomic databases, aiding functional genomics research.
Area of Science:
- Genomics
- Bioinformatics
- Structural Biology
Background:
- Protein family classification is crucial for understanding protein function, evolution, and comparative genomics.
- Existing clustering algorithms face challenges with multi-domain, promiscuous, or fragmented proteins.
Purpose of the Study:
- To introduce TRIBE-MCL, a novel algorithm for rapid and accurate protein sequence clustering into families.
- To address limitations of current methods in handling complex protein structures and large datasets.
Main Methods:
- Utilized the Markov cluster (MCL) algorithm for protein family assignment.
- Employed precomputed sequence similarity information as input for the MCL algorithm.
- Developed TRIBE-MCL to overcome common issues in protein sequence clustering.
Main Results:
- TRIBE-MCL demonstrated rapid and accurate clustering of protein sequences.
- The method successfully handled multi-domain proteins, promiscuous domains, and fragmented proteins.
- Validated on large databases including SwissProt, InterPro, SCOP, and the human genome.
Conclusions:
- TRIBE-MCL is well-suited for large-scale, efficient protein family detection.
- The algorithm facilitates the annotation of a significant portion of proteins in genomic datasets.
- TRIBE-MCL advances functional genomics by enabling robust protein family categorization.