Related Experiment Video
Updated: Sep 7, 2025

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
Recoding Amino Acids to a Reduced Alphabet may Increase or Decrease Phylogenetic Accuracy
Peter G Foster1, Dominik Schrempf2, Gergely J Szöllősi2,3,4
1Department of Life Sciences, Natural History Museum, London SW7 5BD, UK.
Recoding amino acid data can improve phylogenetic accuracy, especially with compositional heterogeneity. However, specific methods like Chi-squared recoding may decrease accuracy, and advanced models offer a more promising solution.
Area of Science:
- Phylogenetics
- Molecular Evolution
- Bioinformatics
Background:
- Amino acid data presents challenges in phylogenetic reconstruction due to long branches and compositional heterogeneity.
- Recoding alignments to reduced alphabets is a common strategy to mitigate these issues.
Purpose of the Study:
- To evaluate the effectiveness of different amino acid recoding strategies on phylogenetic topological accuracy using simulated data.
- To compare recoding methods based on amino acid exchangeability versus those reducing compositional heterogeneity.
Main Methods:
- Simulated four-taxon tree alignments with varying degrees of branch length and compositional heterogeneity.
- Tested three exchangeability-based recoding methods and one Chi-squared statistic-based recoding method.
- Analyzed accuracy using homogeneous phylogenetic models and compared with advanced models (NDCH, CAT).
Main Results:
- Exchangeability-based recoding generally improved accuracy on compositionally heterogeneous alignments.
- Chi-squared recoding showed variable results, sometimes decreasing accuracy.
- Advanced models (NDCH, CAT) analyzing unrecoded, heterogeneous data often outperformed recoded data with homogeneous models.
Conclusions:
- Amino acid recoding can be beneficial but requires cautious interpretation.
- The effectiveness of recoding depends on the type of heterogeneity and the recoding scheme.
- Developing and utilizing better-fitting models (e.g., NDCH, CAT) is a more robust approach for analyzing complex molecular data.
Related Concept Videos
Phylogenetic Trees
Evolutionary Relationships through Genome Comparisons
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Improving Translational Accuracy
From DNA to Protein
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...

