Related Experiment Video
Updated: May 25, 2025

Investigating Protein Sequence-structure-dynamics Relationships with Bio3D-web
Published on: July 16, 2017
20D-Dynamic Representation of Protein Sequences Combined with K-means Clustering
Dorota Bielińska-Wąż1, Piotr Wąż2, Agata Błaczkowska1
1Department of Radiological Informatics and Statistics, Medical University of Gdańsk, 80-210 Gdańsk, Poland.
Objective:
The objective of this research is to demonstrate that alignment-free bioinformatics approaches are effective tools for analyzing the similarity and dissimilarity of protein sequences. All numerical parameters representing sequences are expressed analytically, ensuring precision, clarity, and efficient processing, even for large datasets and long sequences. Additionally, a novel approach for identifying previously unknown virus strains is introduced.
Methods:
A novel approach is proposed, integrating the unique features of our newly developed method, the 20D-Dynamic Representation of Protein Sequences, with the K-means clustering algorithm. The sequences are represented as clouds of material points in a 20-dimensional space (20D-dynamic graphs), with their spatial distribution being unique to each protein sequence. The numerical parameters, referred to as descriptors in molecular similarity theory, represent quantities characteristic of dynamic systems and serve as input data for the K-means clustering algorithm.
Results:
Examples of the application of the approach are presented, including projections of the 20D-dynamic graphs onto 3D spaces, which serve as a visual tool for comparing sequences. Additionally, cluster plots for the analyzed sequences are provided using the proposed method.
Discussion:
Combining the 20D-Dynamic Representation of Protein Sequences with an unsupervised machine learning algorithm (K-means clustering) enhances its scalability. This approach is applicable to large datasets without restrictions on sequence length.
Conclusion:
It has been demonstrated that the 20D-Dynamic Representation of Protein Sequences, combined with the K-means clustering algorithm, successfully classifies subtypes of influenza A virus strains.
More Related Videos
09:17Structure-Based Simulation and Sampling of Transcription Factor Protein Movements along DNA from Atomic-Scale Stepping to Coarse-Grained Diffusion
Published on: March 1, 2022
07:28JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Conservation of Protein Domains
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Protein Dynamics in Living Cells
Fluorescent recovery after photobleaching (FRAP) is a fluorescent-protein-based detection technique used to quantify protein movement rates within the cell. This method exposes a small portion of the cell to an intense laser beam. The laser beam causes permanent photobleaching of the fluorophore-tagged proteins in the exposed region. As the bleached...
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...