Related Experiment Video
Updated: Feb 22, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
Leveraging hierarchical population structure in discrete association studies
Jonathan Carlson1, Carl Kadie, Simon Mallal
1Machine Learning and Applied Statistics Group, Microsoft Research, Redmond, Washington, United States of America; Department of Computer Science and Engineering, University of Washington, Seattle, Washington, United States of America.
Population structure can obscure biological data correlations. This study identifies two confounding processes, coevolution and conditional influence, offering generative models to correct these effects across various biological applications.
Area of Science:
- Genomics
- Population Genetics
- Bioinformatics
Background:
- Population structure is a significant confounder in biological data analysis, leading to spurious correlations.
- Existing methods for correcting population structure effects are disparate and discipline-specific.
- Understanding the distinct processes underlying confounding is crucial for accurate biological inference.
Purpose of the Study:
- To examine methods for correcting confounding in discrete data with hierarchical population structure.
- To identify and define two distinct confounding processes: coevolution and conditional influence.
- To develop and apply generative models for correcting these confounding effects in biological data.
Main Methods:
- Development of generative models to describe coevolution and conditional influence processes.
- Application of these models to correct for confounding in biological datasets.
- Evaluation of model performance across diverse biological applications.
Main Results:
- Identified and characterized two distinct confounding processes: coevolution and conditional influence.
- Demonstrated the utility of generative models in correcting for population structure confounding.
- Showcased successful applications in HIV-1 escape mutation identification, peptide coevolution prediction, and Arabidopsis thaliana resistance trait association.
Conclusions:
- No single method is universally optimal for correcting all forms of population structure confounding.
- The choice between coevolution and conditional influence models depends on the specific biological application.
- Generative models provide a flexible framework for addressing confounding in population genetics and related fields.
More Related Videos
06:52Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
Published on: September 17, 2019
13:55Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization
Published on: February 3, 2013
Related Concept Videos
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Analysis of Population Pharmacokinetic Data
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Confounding in Epidemiological Studies