Related Experiment Video
Updated: Jan 17, 2026

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.7K
Boosting the Power of Rare Variant Association Studies by Imputation Using Large-scale Sequencing Population
Jinglan Dai1, Yixin Zhang1, Yuan Gao1
1Department of Biostatistics, Center for Global Health, School of Public Health, Nanjing Medical University, Nanjing 211166, China.
Genomics, Proteomics & Bioinformatics
|September 19, 2025
Summary
Imputed genetic data, particularly from TOPMed, closely approximates whole-genome sequencing (WGS) for rare variant association studies. Larger sample sizes with imputed data significantly increase the power to detect rare variants, aiding complex trait heritability research.
Area of Science:
- Genomics
- Population Genetics
- Statistical Genetics
Background:
- Population-scale whole-genome sequencing (WGS) enables precise capture of rare variants, crucial for understanding complex trait heritability often missed by conventional genome-wide association studies (GWASs).
- The utility of imputed genetic data versus WGS for rare variant association studies remains an open question.
Purpose of the Study:
- To evaluate the consistency and performance of imputed SNP array data against WGS for rare variant detection.
- To assess the power of imputed data for association studies across various traits and sample sizes.
Main Methods:
- Utilized WGS data (n=150,119) from the UK Biobank as ground truth.
- Assessed imputation quality (TOPMed, HRC+UK10K) for rare variants using R-square and Cramer's V.
- Performed association tests on 45 traits comparing WGS and imputed data at different sample sizes.
- Conducted meta-analysis for lung and ovarian cancer using SNP array and WGS data.
Main Results:
- TOPMed imputation achieved good quality (R-square > 0.6) even for extremely rare variants (minor allele count ≤ 5).
- TOPMed-imputed data showed higher consistency with WGS (Cramer's V > 0.75) across ethnicities.
- At equal sample sizes, imputed data did not outperform WGS, but TOPMed-imputed data results were more concordant.
- Increasing sample size to 488,377 with TOPMed-imputed data boosted rare variant detection by 27.71% (quantitative) and 10-fold (binary).
- Meta-analysis identified more variants and genes compared to WGS-only results.
Conclusions:
- Imputed rare variants from large sequencing cohorts can enhance the power of association tests, especially when WGS sample sizes are limited.
- TOPMed-imputed data offers a valuable resource for rare variant association studies, approaching WGS quality and improving discovery power with scale.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K
RNA-seq
11.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.8K
Next-generation Sequencing
97.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.8K
Sanger Sequencing
773.3K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
773.3K

