HaploBlocks: Efficient Detection of Positive Selection in Large Population Genomic Datasets
Benedikt Kirsch-Gerweck1, Leonard Bohnenkämper2, Michel T Henrichs2
1Palaeogenetics Group, Institute of Organismic and Molecular Evolution (iomE), Johannes Gutenberg University, 55128 Mainz, Germany.
Molecular Biology and Evolution
|February 15, 2023
Summary
We developed an efficient method to detect positive selection in large genomic datasets. This approach is sensitive, specific, and scalable for big data genomics.
Area of Science:
- Genomics
- Population Genetics
- Bioinformatics
Background:
- Genomic regions under positive selection are crucial for understanding adaptation.
- Existing methods for detecting positive selection are computationally intensive, limiting their use on large population genomic datasets.
Purpose of the Study:
- To develop an efficient and scalable haplotype-based method for detecting positive selection in large-scale genomic data.
- To address the computational limitations of current positive selection detection tools.
Main Methods:
- Implemented an efficient haplotype-based approach combining pattern matching (positional Burrows-Wheeler transform) with model-based inference.
- Utilized closed-form expressions for statistical inference to reduce computational load.
- Evaluated the method's performance using simulations and UK Biobank data.
Main Results:
- The developed approach is both sensitive and specific in detecting positive selection.
- Computational resource requirements are low, demonstrating scalability to millions of individuals.
- The method is practical for analyzing large population genomic datasets.
Conclusions:
- The new haplotype-based method provides an efficient and accurate way to detect positive selection.
- This approach is suitable for the demands of "big data" genomics.
- It offers a scalable algorithmic blueprint for future population genomic studies.


