Related Experiment Video
Updated: Jan 3, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
DeepMSA: constructing deep multiple sequence alignment to improve contact prediction and fold-recognition for
Chengxin Zhang1, Wei Zheng1, S M Mortuza1
1Department of Computational Medicine and Bioinformatics, University of Michigan, Ann Arbor, MI 48109, USA.
DeepMSA is a novel method for sensitive multiple sequence alignment (MSA) construction, improving protein structure and function prediction accuracy. This open-source tool enhances bioinformatics applications, especially for proteins lacking known templates.
Area of Science:
- Bioinformatics
- Computational Biology
- Structural Bioinformatics
Background:
- Genome sequencing has led to a vast increase in protein sequences.
- Accurate multiple sequence alignment (MSA) is crucial for modeling protein structure and function.
- Existing pipelines lack efficiency and sensitivity, especially for large-scale genomic and metagenomic data.
Purpose of the Study:
- To develop a novel, sensitive, and efficient pipeline for multiple sequence alignment (MSA) construction.
- To improve the accuracy of protein structure and function prediction methods.
- To provide a robust tool for protein structural bioinformatics applications.
Main Methods:
- Developed DeepMSA, an open-source method for sensitive MSA construction.
- Utilized complementary hidden Markov model algorithms with multi-source databases.
- Evaluated DeepMSA in three large-scale benchmark experiments using 614 non-redundant proteins.
Main Results:
- DeepMSA improved residue-level contact prediction accuracy by up to 24.4% for long-range contacts.
- Average TM-score for homologous structure identification increased by over 7.5%.
- Achieved statistically significant improvements in secondary structure prediction (Q3 accuracy).
Conclusions:
- DeepMSA offers robust and general usefulness in protein structural bioinformatics.
- Improvements were achieved without re-training existing models, demonstrating method's versatility.
- Particularly beneficial for proteins without homologous templates in the Protein Data Bank (PDB).
More Related Videos
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Folding
Protein Folding
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Conservation of Protein Domains
Protein Families