MUMMALS: multiple sequence alignment improved by using hidden Markov models with local structural information

Jimin Pei1, Nick V Grishin

  • 1Howard Hughes Medical Institute, University of Texas Southwestern Medical Center at Dallas, 5323 Harry Hines Boulevard, Dallas, TX 75390-9050, USA. jpei@chop.swmed.edu

Nucleic Acids Research
|August 29, 2006
PubMed

Insights

We developed MUMMALS, a new program for multiple protein sequence alignment. It uses probabilistic consistency and hidden Markov models (HMMs) to improve alignment accuracy, especially for remote homologs.

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Structural Bioinformatics

Background:

  • Accurate multiple protein sequence alignment is crucial for understanding protein function and evolution.
  • Existing alignment methods face challenges with distant protein homologs.

Purpose of the Study:

  • To develop a novel program, MUMMALS, for enhanced multiple protein sequence alignment.
  • To improve alignment accuracy by incorporating probabilistic consistency and local structural information.

Main Methods:

  • Developed MUMMALS, a program utilizing probabilistic consistency for multiple sequence alignment.
  • Employed pairwise alignment hidden Markov models (HMMs) with multiple match states.
  • Estimated model parameters using a large library of structure-based alignments.

Main Results:

  • MUMMALS demonstrated statistically superior accuracy compared to leading aligners (ProbCons, MAFFT, MUSCLE) on remote homologs.
  • A large dataset of automatically computed pairwise structure alignments proved more effective for parameter estimation and testing than smaller curated datasets.
  • Reference-independent evaluation methods showed strong correlation with reference-dependent evaluations.

Conclusions:

  • MUMMALS offers improved accuracy for multiple protein sequence alignment, particularly for distantly related proteins.
  • Large-scale, automatically generated datasets are valuable for training and evaluating alignment algorithms.
  • Validated a reliable method for reference-independent assessment of alignment quality.

Related Concept Videos

Multi-species Conserved Sequences02:51

Multi-species Conserved Sequences

Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale  studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Modern Molecular Taxonomy01:29

Modern Molecular Taxonomy

Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...