Related Experiment Video
Updated: Oct 10, 2025

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
A Fundamental Problem with Amino-Acid-Sequence Characters for Phylogenetic Analyses
1L. H. Bailey Hortorium, Cornell University, 462 Mann Library, Ithaca, New York, 14853.
Abstract:
Protein-coding genes may be analyzed in phylogenetic analyses using nucleotide-sequence characters and/or amino-acid-sequence characters. Although amino-acid-sequence characters "correct" for saturation (parallelism), amino-acid-sequence characters are subject to convergence and ignore phylogenetically informative variation. When all nucleotide-sequence characters have a consistency index of 1, characters coded using the amino acid sequence may have a consistency index of less than 1. The reason for this is that most amino acids are specified by more than one codon. If two different codons that both code for the same amino acid are derived independent of one another in divergent lineages, nucleotide-sequence characters may not be homoplasious when amino-acid-sequence characters may be homoplasious. Not only may amino-acid-sequence characters support groupings that are not supported by nucleotide-sequence characters, they may support contradictory groupings. Because this convergence is a problem of character delimitation, it affects the results of all tree-construction methods (maximum likelihood, neighbor joining, parsimony, etc.). In effect, coding amino-acid-sequence characters instead of nucleotide-sequence characters putatively corrects for saturation and definitely causes a convergence problem. An empirical example from the Mhc locus is given.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Phylogenetic Trees
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Amino acids
Phylogeny
Protein Families

