Related Experiment Video
Updated: Jul 20, 2026

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
MUMMALS: multiple sequence alignment improved by using hidden Markov models with local structural information
1Howard Hughes Medical Institute, University of Texas Southwestern Medical Center at Dallas, 5323 Harry Hines Boulevard, Dallas, TX 75390-9050, USA. jpei@chop.swmed.edu
Abstract:
We have developed MUMMALS, a program to construct multiple protein sequence alignment using probabilistic consistency. MUMMALS improves alignment quality by using pairwise alignment hidden Markov models (HMMs) with multiple match states that describe local structural information without exploiting explicit structure predictions. Parameters for such models have been estimated from a large library of structure-based alignments. We show that (i) on remote homologs, MUMMALS achieves statistically best accuracy among several leading aligners, such as ProbCons, MAFFT and MUSCLE, albeit the average improvement is small, in the order of several percent; (ii) a large collection (>10 000) of automatically computed pairwise structure alignments of divergent protein domains is superior to smaller but carefully curated datasets for estimation of alignment parameters and performance tests; (iii) reference-independent evaluation of alignment quality using sequence alignment-dependent structure superpositions correlates well with reference-dependent evaluation that compares sequence-based alignments to structure-based reference alignments.
Insights
We developed MUMMALS, a new program for multiple protein sequence alignment. It uses probabilistic consistency and hidden Markov models (HMMs) to improve alignment accuracy, especially for remote homologs.
Area of Science:
- Bioinformatics
- Computational Biology
- Structural Bioinformatics
Background:
- Accurate multiple protein sequence alignment is crucial for understanding protein function and evolution.
- Existing alignment methods face challenges with distant protein homologs.
Purpose of the Study:
- To develop a novel program, MUMMALS, for enhanced multiple protein sequence alignment.
- To improve alignment accuracy by incorporating probabilistic consistency and local structural information.
Main Methods:
- Developed MUMMALS, a program utilizing probabilistic consistency for multiple sequence alignment.
- Employed pairwise alignment hidden Markov models (HMMs) with multiple match states.
- Estimated model parameters using a large library of structure-based alignments.
Main Results:
- MUMMALS demonstrated statistically superior accuracy compared to leading aligners (ProbCons, MAFFT, MUSCLE) on remote homologs.
- A large dataset of automatically computed pairwise structure alignments proved more effective for parameter estimation and testing than smaller curated datasets.
- Reference-independent evaluation methods showed strong correlation with reference-dependent evaluations.
Conclusions:
- MUMMALS offers improved accuracy for multiple protein sequence alignment, particularly for distantly related proteins.
- Large-scale, automatically generated datasets are valuable for training and evaluating alignment algorithms.
- Validated a reliable method for reference-independent assessment of alignment quality.
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...

