Related Experiment Video
Updated: Aug 9, 2025

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Clustering Highly Divergent Homologous Proteins: An Alignment-Free Method
Laura Muñoz-Baena1, Art F Y Poon1,2
1Department of Microbiology and Immunology, Western University, London, Ontario, Canada.
This study introduces an alignment-free method to classify homologous protein-coding regions across different genomes. It uses k-mer frequency distributions and clustering to identify related sequences, improving genomic comparisons.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Comparative analysis of amino acid sequences is crucial in molecular biology.
- Accurate alignment of protein-coding sequences becomes challenging for distantly related genomes.
- Identifying homologous regions across different genomes is difficult with traditional methods.
Purpose of the Study:
- To present an alignment-free method for classifying homologous protein-coding regions from diverse genomes.
- To adapt a methodology originally for virus families to broader genomic comparisons.
- To quantify sequence homology using k-mer frequency distributions and intersection distance.
Main Methods:
- Quantifying sequence homology via the overlap of k-mer frequency distributions (intersection distance).
- Employing dimensionality reduction and hierarchical clustering to extract homologous sequence groups from a distance matrix.
- Generating visualizations of cluster composition and genome-wide cluster assignments for reliability assessment.
Main Results:
- Demonstrated an effective alignment-free approach for identifying homologous protein-coding regions.
- Successfully applied k-mer distance, dimensionality reduction, and hierarchical clustering for sequence classification.
- Developed visualization tools to assess the reliability and biological relevance of clustering results.
Conclusions:
- The proposed alignment-free method offers a robust alternative for classifying homologous protein-coding regions, especially in challenging cross-genome comparisons.
- The methodology is adaptable beyond virus families to other organisms.
- Visualizations aid in the rapid evaluation of clustering accuracy and biological interpretation.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Families
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Conservation of Protein Domains
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...

