Related Experiment Video
Updated: Dec 22, 2025

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
Ambiguity Coding Allows Accurate Inference of Evolutionary Parameters from Alignments in an Aggregated State-Space
Claudia C Weber1, Umberto Perron1, Dearbhaile Casey1
1European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton CB10 1SD, UK.
Learn protein evolution history by adapting models for missing data. This method recovers evolutionary information from sequences previously inaccessible, improving ancestral reconstruction and selection strength estimates.
Area of Science:
- Evolutionary Biology
- Computational Biology
- Biophysics
Background:
- Understanding protein evolution requires models that capture genetic variation and functional constraints.
- Practical approaches often balance data availability with useful parameter estimation, such as selection strength or ancestral structure.
- Limited data resolution can hinder the application of advanced evolutionary models.
Purpose of the Study:
- To demonstrate a method for obtaining accurate evolutionary parameter estimates from data with limited resolution.
- To show how to infer ancestral protein states and evolutionary parameters when complete data is unavailable.
- To improve ancestral reconstruction and the estimation of selection strength using adapted substitution models.
Main Methods:
- Encoding observed characters as ambiguous representations within a larger state-space.
- Applying established methods for handling missing data to protein sequence alignments.
- Utilizing adapted codon models (e.g., 61-state) and empirical models (e.g., 55-state) for amino acid data.
Main Results:
- Accurate and unbiased estimation of the selection strength parameter omega ($\omega$) using an adapted 61-state codon model.
- Successful inference of ancestral amino acid side chain configurations using a 55-state model on 20-state amino acid data.
- Significantly improved ancestral reconstruction accuracy by incorporating structural information into even a small fraction of sequences.
Conclusions:
- A novel strategy allows the recovery of crucial evolutionary information from protein sequences with previously inaccessible data.
- This ambiguity-coding approach expands the applicability of sophisticated evolutionary models to lower-resolution sequence data.
- The methods presented enhance our ability to reconstruct protein evolutionary history and understand natural selection.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Gene Evolution - Fast or Slow?
Synteny and Evolution
Around 80 million years ago, the human and mice lineages diverged from the common ancestor. During the course of evolution, the ancestral...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Cis-regulatory Sequences

