Related Experiment Video
Updated: Dec 19, 2025

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Obtaining extremely large and accurate protein multiple sequence alignments from curated hierarchical alignments
Andrew F Neuwald1,2, Christopher J Lanczycki3, Theresa K Hodges1
1Institute for Genome Sciences.
This study introduces hierarchical multiple sequence alignments (hiMSAs) to improve protein sequence analysis. These curated alignments, along with the MAPGAPS program, enable accurate analysis of large sequence datasets for machine learning applications.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning in genomics
Background:
- Machine learning for protein analysis requires large multiple sequence alignments (MSAs).
- Existing automated MSA generation methods (e.g., PSI-BLAST, JackHMMER) struggle with distant homologs and evolutionary divergence, necessitating manual curation.
- Manually curated MSAs are often too small for statistical methods.
Purpose of the Study:
- To address limitations in automated MSA generation for machine learning.
- To provide a robust method for creating large, accurate MSAs from curated hierarchical MSAs (hiMSAs).
- To enable improved protein sequence and structural analysis using deep learning and big data approaches.
Main Methods:
- Development and release of 252 curated hiMSAs comprising over 26 million sequences.
- Introduction of the MAPGAPS search program for querying hiMSAs to align vast numbers of database sequences.
- Hierarchical MSA construction with subgroup and template alignments for accurate sequence relationships.
Main Results:
- Demonstrated accurate alignment of database sequences using hiMSAs and MAPGAPS, comparable to curated MSA quality.
- Successfully applied the method to the exonuclease-endonuclease-phosphatase superfamily and pleckstrin homology domains.
- Generated extremely large MSAs suitable for deep learning and big data analyses.
Conclusions:
- The developed hiMSA framework and MAPGAPS program significantly enhance the accuracy and scale of MSAs for computational biology.
- This resource facilitates advanced machine learning applications in protein sequence and structural analysis.
- Public availability of hiMSAs and associated software promotes further research in the field.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Conservation of Protein Domains
Genome Annotation and Assembly
Protein Organization
The primary structure of a protein is its amino acid sequence....

