Related Experiment Video
Updated: Jul 13, 2026

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
On reduced amino acid alphabets for phylogenetic inference
1Department of Mathematics and Statistics, Dalhousie University, Halifax, Nova Scotia, Canada. susko@mathstat.dal.ca
Molecular Biology and Evolution
|July 27, 2007
Summary
Using reduced amino acid alphabets with Markov models of evolution minimizes errors in evolutionary studies. This binning approach improves accuracy, especially with complex biological data.
Area of Science:
- Evolutionary biology
- Computational biology
- Bioinformatics
Background:
- Markov models of evolution are crucial for phylogenetic analysis.
- Model misspecification and saturation can bias evolutionary inferences.
- Reduced amino acid alphabets offer a potential solution to these issues.
Purpose of the Study:
- To investigate the effectiveness of Markov models using reduced amino acid alphabets (bins).
- To develop and evaluate algorithms for automated bin construction.
- To assess the impact of binning on phylogenetic tree estimation.
Main Methods:
- Development of algorithms for binning amino acids based on rate matrices and alignment properties.
- Simulations to compare binning approaches with standard methods and missing data approaches.
- Application of binning methods to real biological datasets exhibiting compositional heterogeneity and saturation.
Main Results:
- Binning amino acids results in minimal information loss, even without model misspecification.
- Markov models utilizing binned amino acids perform comparably to more complex missing data methods.
- Binning demonstrably improves topological accuracy in phylogenetic estimation for real datasets.
Conclusions:
- Reduced amino acid alphabets provide a practical strategy to mitigate model misspecification and saturation in evolutionary analyses.
- Automated binning algorithms offer an efficient way to implement this approach.
- This method enhances the reliability of phylogenetic tree reconstruction in challenging biological scenarios.
Related Concept Videos
Phylogenetic Trees
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.The length of the branches can depict time or the relative amount of change among organisms. For instance, the branch length might indicate the number of amino acid changes in the sequence that underlies the...
Phylogenetic Trees
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.The length of the branches can depict time or the relative amount of change among organisms. For instance, the branch length might indicate the number of amino acid changes in the sequence that underlies the...
Amino acids
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible for...
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...

