Related Experiment Video
Updated: Jun 15, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
A rapid method for characterization of protein relatedness using feature vectors
Kareem Carr1, Eleanor Murray, Ebenezer Armah
1Department of Mathematics, Statistics and Computer Science, University of Illinois at Chicago, Chicago, Illinois, United States of America.
This study introduces a novel feature vector method for analyzing biological sequence variations without alignment. This approach effectively visualizes protein relationships and evolutionary patterns, aiding in phylogenetic and functional analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Molecular Evolution
Background:
- Analyzing large biological sequence datasets is crucial for understanding molecular evolution and protein function.
- Traditional methods like sequence alignment can be computationally intensive and may introduce biases.
- A need exists for alignment-free methods to explore relationships within protein families.
Purpose of the Study:
- To develop and validate a novel feature vector approach for characterizing variations in large biological sequence datasets.
- To enable visualization of protein relationships without relying on sequence alignment or prior assumptions.
- To assess the utility of this method for identifying phylogenetic and functional relationships.
Main Methods:
- Constructing feature vectors based on the count and location of amino acids or nucleic acids in each sequence.
- Utilizing binomial and uniform distributions to model theoretical sequences and calculate distances.
- Applying principal component analysis (PCA) for high-dimensional feature vector interpretation and visualization.
- Employing agglomerative hierarchical clustering to construct cladograms from feature vectors.
Main Results:
- The feature vector approach effectively captures significant information from the original biological sequences.
- PCA successfully extracts key variation patterns from collections of feature vectors, enabling rapid analysis.
- The method accurately identifies phylogenetically and functionally related protein collections, including protein kinase C and globins.
- Hierarchical clustering generates meaningful cladograms from the feature vector data.
Conclusions:
- The proposed feature vector method offers an efficient and alignment-free alternative for analyzing biological sequence variation.
- This approach facilitates the visualization and interpretation of complex relationships within protein families.
- The method holds promise for advancing studies in molecular evolution, protein function, and phylogenetic analysis.
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Evolutionary Relationships through Genome Comparisons
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Modern Molecular Taxonomy
Two-dimensional Gel Electrophoresis
The first dimension separation uses the isoelectric focusing or IEF technique performed on immobilized pH gradient (IPG) strips that separate proteins according to their isoelectric points.
Biological samples, such as cells...

