Related Experiment Video
Updated: Jul 7, 2025

04:52
Following the Dynamics of Structural Variants in Experimentally Evolved Populations
Published on: February 3, 2023
990
Simulations of Sequence Evolution: How (Un)realistic They Are and Why.
Johanna Trost1, Julia Haag2, Dimitri Höhler2
1Biometry and Evolutionary Biology Laboratory (LBBE), University Claude Bernard Lyon 1, Lyon, France.
Molecular Biology and Evolution
|December 21, 2023
Summary
Current models struggle to simulate realistic multiple sequence alignments (MSAs). Machine learning can distinguish simulated from empirical MSAs, revealing limitations in evolutionary models.
Area of Science:
- Bioinformatics
- Computational Biology
- Evolutionary Biology
Background:
- Probabilistic models of sequence evolution are vital for evaluating phylogenetic inference tools and developing machine learning (ML) approaches for phylogenetic reconstruction.
- Realistic simulation of multiple sequence alignments (MSAs) is crucial for assessing the performance of phylogenetic tools on empirical data and ensuring ML models generalize well.
Purpose of the Study:
- To simulate DNA and protein MSAs using state-of-the-art methods under various evolutionary models, with and without insertion/deletion (indel) events.
- To assess the realism of simulated MSAs by quantifying the accuracy of supervised learning methods in distinguishing them from empirical MSAs.
Main Methods:
- Simulation of DNA and protein MSAs using a state-of-the-art sequence simulator under progressively complex evolutionary models.
- Application of two distinct, independently developed supervised learning classification approaches to differentiate simulated from empirical MSAs.
Main Results:
- High accuracy was achieved in distinguishing between empirical and simulated MSAs across all tested evolutionary models using both classification methods.
- Findings indicate that current evolutionary models inadequately capture key characteristics of empirical MSAs, such as site-wise evolutionary rates and nucleotide/amino acid composition.
Conclusions:
- Existing probabilistic models of sequence evolution do not fully replicate the complexity of empirical sequence data.
- Supervised learning methods can effectively identify subtle differences between simulated and empirical MSAs, highlighting areas for improvement in evolutionary modeling.
Related Concept Videos
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
The Evidence for Evolution
42.8K
Genetic variations accumulating within populations over generations give rise to biological evolution. Evolutionary changes can result in the formation of novel varieties and entire new species. These changes are responsible for the diverse forms of life inhabiting the planet. The evidence for evolution suggests that all living organisms descended from common ancestors.
42.8K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Hardy-Weinberg Principle
72.2K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
72.2K
Maxam-Gilbert Sequencing
11.2K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.2K

