Related Experiment Video
Updated: Apr 21, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Context similarity scoring improves protein sequence alignments in the midnight zone.
Armin Meier1, Johannes Söding2
1Gene Center, LMU Munich, 81377 Munich and Max Planck Institute for Biophysical Chemistry, 37077 Göttingen, Germany.
This study introduces an unsupervised method to learn conserved sequence patterns, enhancing protein alignment accuracy. The new context similarity score significantly improves protein structure prediction and homology modeling.
Area of Science:
- Bioinformatics
- Computational Biology
- Structural Bioinformatics
Background:
- High-quality protein sequence alignments are critical for applications like template-based protein structure prediction.
- Current profile-profile alignment tools often incorporate 1D structural properties (e.g., secondary structure, solvent accessibility) predicted from local sequence windows.
- These additional terms evaluate the conservation of local amino acid properties, capturing correlations between profile columns.
Purpose of the Study:
- To develop a novel, agnostic approach for learning conserved patterns in protein sequences.
- To improve the accuracy of protein sequence alignments by incorporating a new context similarity score.
- To enhance downstream applications reliant on precise pairwise alignments.
Main Methods:
- An unsupervised learning approach was employed to identify maximally conserved patterns.
- A set of 32 conserved patterns, represented by 13-residue sequence profiles, were trained using a maximum likelihood approach.
- These patterns were learned from structurally aligned training profiles of remotely homologous protein pairs.
- The developed context score was integrated into the Hmm-Hmm alignment tool (hhsearch).
Main Results:
- The study successfully learned a set of 32 conserved sequence profiles in an unsupervised manner.
- The integration of the new context score into hhsearch led to significant improvements in the quality of difficult protein sequence alignments.
- The agnostic approach avoids the need to pre-define or predict specific 1D structural properties.
Conclusions:
- The context similarity score demonstrably enhances the quality of protein homology models.
- This improvement extends to other methods that rely on accurate pairwise sequence alignments.
- The findings highlight the utility of learning conserved sequence patterns directly from data for improved alignment accuracy.
More Related Videos
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Complex Assembly
Protein Complex Assembly
Many viruses self-assemble into a fully functional unit using the infected host cell to...
Conservation of Protein Domains

