Related Experiment Video
Updated: May 11, 2026

07:49
Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Top-down clustering for protein subfamily identification
Eduardo P Costa1, Celine Vens, Hendrik Blockeel
1Department of Computer Science, KU Leuven, Belgium.
Evolutionary Bioinformatics Online
|May 24, 2013
Summary
We developed a new method for protein subfamily identification that builds a hierarchical tree top-down. This approach improves cluster accuracy and identifies key mutations for classifying new protein sequences.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Protein families contain diverse sequences with varying functions.
- Accurate identification of protein subfamilies is crucial for understanding protein function and evolution.
- Existing methods for subfamily identification have limitations in accuracy and interpretability.
Purpose of the Study:
- To propose a novel top-down hierarchical tree construction method for protein subfamily identification.
- To associate specific mutations with protein subclusters for functional insights.
- To improve the accuracy of subfamily identification and facilitate the classification of new protein sequences.
Main Methods:
- Building a hierarchical tree from multiple protein sequence alignments using a top-down approach.
- Employing a post-pruning procedure to extract clusters from the constructed tree.
- Associating specific mutations with each subcluster division.
Main Results:
- The novel method yields more accurate clusters and a better tree topology compared to the state-of-the-art SCI-PHY method.
- The method successfully identifies known functional sites within protein families.
- Specific mutations identified by the method allow for accurate classification of new protein sequences, approaching hidden Markov model accuracy.
Conclusions:
- The proposed top-down hierarchical method enhances protein subfamily identification accuracy.
- The method provides valuable insights into functionally important sites and mutations.
- This approach offers a more effective tool for classifying novel protein sequences.
Related Concept Videos
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...

