Related Experiment Video
Updated: May 10, 2025

09:34
Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.4K
Genetic Similarity Clustering Using the UK Biobank as a Reference Dataset
Ngoc-Quynh Le1,2, Puya Gharahkhani1, Stuart MacGregor1
1Statistical Genetics Lab, QIMR Berghofer Medical Research Institute, Herston, Brisbane, QLD, Australia.
Summary
This study used UK Biobank data to cluster genetically similar individuals, creating a more diverse reference panel. This approach enhances genetic studies for underrepresented populations, improving health equity.
Area of Science:
- Genomics
- Population Genetics
- Bioinformatics
Background:
- Diverse genetic data is vital for disease research and health equity.
- Current reference panels lack population diversity and sufficient sample sizes.
- UK Biobank offers a large dataset for improving population representation.
Purpose of the Study:
- To evaluate the UK Biobank as a reference for clustering genetically similar individuals.
- To enhance population representation and mitigate bias in genetic studies.
- To develop a practical method for inferring population labels using biobank data.
Main Methods:
- Combined UK Biobank country of birth and ethnic background data with genetic information.
- Utilized a random forest model trained on genetic principal components to assign population labels.
- Validated the model using 1000 Genomes and CARTaGENE biobank data.
Main Results:
- Identified 19 diverse reference populations globally, exceeding the diversity of existing datasets like 1000 Genomes.
- Achieved medium to high precision and recall for most identified populations.
- Demonstrated high precision (81.1%) and recall (97.0%) for a Middle Eastern reference sample in CARTaGENE data.
Conclusions:
- Clustering genetically similar individuals using UK Biobank data creates a more representative reference panel.
- This approach can facilitate downstream genetic analyses, including genomewide association studies and polygenic risk scores.
- The method supports research in underrepresented populations, promoting health equity.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
5.6K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.6K
Genome-wide Association Studies-GWAS
12.0K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.0K
Gene Evolution - Fast or Slow?
7.0K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.0K

