Related Experiment Video
Updated: Jul 16, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Scaling the discrete-time Wright-Fisher model to biobank-scale datasets
Jeffrey P Spence1, Tony Zeng1, Hakhamanesh Mostafavi1
1Department of Genetics, Stanford University, Stanford, CA 94305, USA.
A new scalable algorithm approximates the discrete-time Wright-Fisher model, enabling population genetics analysis for millions. This method improves the accuracy of estimating selection coefficients from large genetic datasets.
Area of Science:
- Population Genetics
- Computational Biology
- Bioinformatics
Background:
- The discrete-time Wright-Fisher (DTWF) model is fundamental for understanding allele frequency evolution due to genetic drift, mutation, and selection.
- Current computational methods for DTWF likelihoods are limited by scalability, failing for large sample sizes common in exome sequencing.
- Diffusion approximations, while computationally tractable, become inaccurate with large samples or strong selection.
Purpose of the Study:
- To develop a scalable algorithm for approximating the DTWF model with provably bounded error.
- To enable accurate population genetics inference for biobank-scale datasets (millions of individuals).
- To assess the impact of increasing sample sizes on estimating selection coefficients, particularly for loss-of-function variants.
Main Methods:
- Leveraged two key DTWF properties: approximate sparsity of transition probabilities and closeness of transition distributions for similar allele frequencies.
- Developed an approximate matrix-vector multiplication technique achieving linear time complexity.
- Extended these principles to Hypergeometric distributions for efficient subsampling likelihood computations.
Main Results:
- The novel algorithm accurately approximates the DTWF model and scales to population sizes in the tens of millions.
- Theoretical and practical demonstrations confirm the high accuracy of the approximation.
- Analysis indicates that sample sizes beyond current large exome sequencing cohorts offer diminishing returns for estimating selection coefficients, except for genes with extreme fitness effects.
Conclusions:
- The developed scalable algorithm overcomes previous computational limitations in population genetics.
- This facilitates rigorous, large-scale inference, including biobank-level analyses.
- Findings suggest current large exome sequencing efforts are near the point of diminishing returns for selection coefficient estimation in most genes.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Biostatistics: Overview
Discrete variables are...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Model Approaches for Pharmacokinetic Data: Physiological Models
BIBO stability of continuous and discrete -time systems
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....

