Related Experiment Video
Updated: Jul 7, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
A study of residue correlation within protein sequences and its application to sequence classification
1Center For Genomics and Bioinformatics, Indiana University, 1001 E. 3rd Street, Bloomington, Indiana, IN 47405-3700, USA.
We developed a new method, the mutual information vector (MIV), to better estimate residue correlations in protein sequences. This approach significantly improves protein family classification without needing sequence alignment.
Area of Science:
- Computational biology
- Bioinformatics
- Protein sequence analysis
Background:
- Estimating residue correlations is crucial for understanding protein structure and function.
- Traditional methods like mutual information (MI) primarily capture local interactions.
- There is a need for methods that can detect long-range correlations in protein sequences.
Purpose of the Study:
- To develop and evaluate novel methods for estimating residue correlations in protein sequences.
- To improve upon existing mutual information (MI) based approaches.
- To assess the utility of these methods for protein classification tasks.
Main Methods:
- Calculating mutual information (MI) for adjacent residues.
- Defining the mutual information vector (MIV) to capture long-range residue correlations.
- Incorporating residue hydropathy as an alternative correlation metric.
- Conducting protein family classification tests using the developed methods.
Main Results:
- The mutual information vector (MIV) demonstrated superior performance compared to the classic MI method.
- MIV-based modeling significantly enhanced the accuracy of protein family classification.
- The method achieved a level of classification performance that did not require sequence alignment.
Conclusions:
- The mutual information vector (MIV) is a powerful tool for estimating residue correlations, including long-range interactions.
- MIV significantly outperforms traditional MI methods in protein family classification.
- This alignment-free approach offers a promising direction for protein sequence analysis and classification.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein Families
Protein Families

