Shrinkage-based Random Local Clocks with Scalable Inference
Alexander A Fisher1, Xiang Ji2, Akihiko Nishimura3
1Department of Statistical Science, Duke University, Durham, NC, USA.
Molecular Biology and Evolution
|November 11, 2023
Summary
We developed a novel Bayesian model for molecular clock evolution, improving divergence-time estimation scalability and accuracy. This shrinkage clock method efficiently handles large phylogenetic trees and complex evolutionary rates.
Area of Science:
- Computational Biology
- Phylogenetics
- Evolutionary Biology
Background:
- Molecular clock models are crucial for estimating divergence times in evolutionary biology.
- Existing local clock models face limitations in scalability, model misspecification, and require prior knowledge of clock locations.
- Efficient and accurate molecular clock inference is essential for understanding evolutionary history.
Purpose of the Study:
- To present a new autocorrelated, Bayesian model for heritable clock rate evolution that overcomes limitations of current methods.
- To develop an efficient computational method for scaling molecular clock inference to large phylogenetic trees.
- To apply the new model to infer evolutionary rates in mammalian and influenza virus phylogenies.
Main Methods:
- Developed an autocorrelated Bayesian model with heavy-tailed priors for heritable clock rate evolution.
- Implemented an efficient Hamiltonian Monte Carlo sampler with closed-form gradient computations for scalability.
- Applied the 'shrinkage clock' model to simulated datasets, mammalian phylogenies, and influenza A virus surface glycoproteins.
Main Results:
- The shrinkage clock model demonstrates significant speed-up compared to random local clock methods, especially for large datasets.
- The model successfully recovers known local clock structures in rodent and mammalian phylogenies.
- Enabled computationally intensive analysis of influenza A virus surface glycoprotein evolution without prior clock placement assumptions.
Conclusions:
- The proposed shrinkage clock model offers a scalable and accurate approach to molecular clock inference.
- This method advances the estimation of divergence times and the understanding of evolutionary rate variation.
- The publicly available implementation in BEAST facilitates broader application in evolutionary and phylogenetic studies.
Related Concept Videos
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Propagation of Uncertainty from Random Error
703
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
703
Randomized Experiments
7.0K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
7.0K
Random Sampling Method
11.2K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
11.2K
Random Variables
12.3K
A random variable is a single numerical value that indicates the outcome of a procedure. The concept of random variables is fundamental to the probability theory and was introduced by a Russian mathematician, Pafnuty Chebyshev, in the mid-nineteenth century.
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
12.3K
Random Error
896
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
896


