Related Experiment Video
Updated: Sep 11, 2025

12:00
A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
35.5K
The Structure of Deviations From Maximum Parsimony for Densely-Sampled Data and Applications for Clade Support
IEEE Transactions on Computational Biology and Bioinformatics
|August 14, 2025
Summary
Maximum parsimony (MP) algorithms can produce incorrect phylogenetic trees, especially with dense viral sampling. This study develops algorithms to understand these deviations, finding they often occur locally due to independent mutations on sister branches.
Area of Science:
- Computational Biology
- Phylogenetics
- Evolutionary Biology
Background:
- Phylogenetic reconstruction algorithms aim to infer evolutionary histories but can produce incorrect trees.
- Maximum parsimony (MP) is a fundamental phylogenetic criterion, yet its failure modes, especially under dense sampling, are not fully understood.
- Dense sampling, common in viral evolution studies like SARS-CoV-2, presents unique challenges for phylogenetic inference.
Purpose of the Study:
- To investigate how phylogenetic reconstruction algorithms, specifically maximum parsimony, deviate from the correct evolutionary tree in densely sampled datasets.
- To understand the structural nature of these deviations in the context of near-maximally parsimonious evolutionary histories.
- To develop novel algorithms for analyzing these deviations and improving phylogenetic inference.
Main Methods:
- Development of new algorithms to analyze the structural differences between correct evolutionary trees and MP trees.
- Application of these algorithms to simulated datasets mimicking SARS-CoV-2 evolution under dense sampling conditions.
- Analysis of the local structures causing deviations from maximal parsimony.
Main Results:
- Deviations from maximally parsimonious trees in densely sampled simulations are frequently local.
- These local deviations often involve the independent appearance of the same mutation on sister branches within the phylogenetic tree.
- The identified patterns provide insights into the structure of errors in MP-based phylogenetics for dense viral data.
Conclusions:
- Understanding local deviations from maximal parsimony is crucial for accurate phylogenetic reconstruction with dense viral sequence data.
- The findings enable the design of methods to sample near-MP trees more effectively.
- This approach facilitates efficient estimation of clade supports, improving the reliability of evolutionary inferences.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
6.1K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.1K
Phylogenetic Trees
46.4K
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.
46.4K
Phylogeny
47.3K
Phylogeny is concerned with the evolutionary diversification of organisms or groups of organisms. A group of organisms with a name is called a taxon (singular). Taxa (plural) can span different levels of the evolutionary hierarchy. For instance, the group containing all birds is a taxon (comprising the class Aves), and the group of all species of daisies (the genus Bellis) is a taxon. Phylogenies can likewise include just one genus (i.e., depict species relationships) or span an entire kingdom.
47.3K
Distributions to Estimate Population Parameter
4.3K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.3K
Estimating Population Standard Deviation
3.1K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.1K
Gene Evolution - Fast or Slow?
7.4K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.4K

