HHalign-Kbest: exploring sub-optimal alignments for remote homology comparative modeling.
Jinchao Yu1, Geraldine Picord2, Pierre Tuffery2
1Institute for Integrative Biology of the Cell (I2BC), CEA, CNRS, University Paris-Saclay, CEA-Saclay, 91191 Gif-sur-Yvette.
Bioinformatics (Oxford, England)
|August 2, 2015
Summary
The HHalign-Kbest server improves protein structure prediction by generating multiple suboptimal alignments, enhancing accuracy in the twilight zone of sequence identity. This method systematically improves model quality compared to using only the optimal alignment.
Area of Science:
- Computational biology
- Bioinformatics
- Structural biology
Background:
- The HHsearch algorithm uses hidden Markov model (HMM)-HMM alignment for protein sequence alignment, performing well even with low sequence identity (twilight zone).
- However, optimal HHsearch alignments can contain errors, negatively impacting protein structure prediction accuracy.
Purpose of the Study:
- To develop a novel algorithm, HHalign-Kbest, for generating k-best suboptimal HMM-HMM alignments within the HHsearch framework.
- To improve protein structure prediction by evaluating multiple suboptimal alignments and selecting the best structural models.
Main Methods:
- Implemented a novel algorithm to generate k-best suboptimal HMM-HMM alignments, moving beyond the single optimal alignment.
- Utilized a directed acyclic graph-based approach for efficient memory usage in large protein alignments.
- Systematically generated and evaluated structural models using the Qmean score for top k suboptimal alignments.
Main Results:
- The HHalign-Kbest server demonstrated improved alignment quality among the top k suboptimal alignments.
- Benchmarking on 420 SCOP30 targets showed significant increases in model quality (TM-score).
- Average model quality improved by 4.1-16.3% (top 1 model) and 8.0-21.0% (top 10 models) for HHsearch probabilities between 20-99%.
Conclusions:
- The HHalign-Kbest server effectively enhances protein structure prediction by leveraging suboptimal alignments.
- This approach offers a systematic way to improve the accuracy of protein structure models, particularly in challenging alignment scenarios.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
7.3K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.3K
Conserved Binding Sites
5.3K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.3K
Conservation of Protein Domains Over Different Proteins
15.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
15.0K


