Related Experiment Video
Updated: Jul 7, 2026

09:51
Investigating Protein Sequence-structure-dynamics Relationships with Bio3D-web
Published on: July 16, 2017
ProClust: improved clustering of protein sequences with an extended graph-based approach
P Pipenbacher1, A Schliep, S Schneckener
1ZAIK/ZPR, Universität zu Köln, Germany.
Bioinformatics (Oxford, England)
|October 19, 2002
Summary
This study introduces an improved graph-based clustering algorithm for identifying remote protein homologues. The method enhances accuracy by considering sequence similarity, alignment score significance, and transitivity, outperforming existing tools like PSI-Blast.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Identifying remote protein homologues is challenging due to increasing database sizes and noise.
- Traditional alignment methods struggle with distinguishing true homologues from random similarities.
- Protein families and transitivity of homology offer potential solutions but have limitations, especially with multi-domain proteins.
Purpose of the Study:
- To develop an improved algorithm for detecting remote protein homologues.
- To enhance the accuracy and sensitivity of homology detection in large biological databases.
- To address limitations of existing methods, particularly concerning multi-domain proteins and noise.
Main Methods:
- An extended graph-based clustering algorithm utilizing an asymmetric distance measure.
- Scaling similarity values by sequence length and incorporating alignment score significance for filtering.
- Post-processing with profile hidden Markov models (HMMs) for cluster merging.
- Utilizing transitivity with multiple intermediate sequences.
Main Results:
- The proposed method demonstrates high specificity and favorably compares with PSI-Blast for remote homologue detection.
- The algorithm effectively uses transitivity, with up to twelve intermediate sequences, crucial for performance.
- Analysis of false positives indicates effective bounding of transitivity degree and provides parameter guidance.
- The approach offers substantial improvement over existing methods, despite theoretical limitations with multi-domain proteins.
Conclusions:
- The developed algorithm significantly improves remote protein homology detection.
- The method effectively leverages transitivity and addresses challenges posed by large datasets and noise.
- While not a complete solution for multi-domain proteins, the heuristics provide a substantial advancement in the field.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...

