Related Experiment Video
Updated: Jan 1, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
UK Biobank Whole-Exome Sequence Binary Phenome Analysis with Robust Region-Based Rare-Variant Test
Zhangchen Zhao1, Wenjian Bi1, Wei Zhou2
1Department of Biostatistics, University of Michigan School of Public Health, Ann Arbor, MI 48109, USA; Center for Statistical Genetics, University of Michigan School of Public Health, Ann Arbor, MI 48109, USA.
New region-based tests accurately analyze unbalanced biobank data, improving genetic association studies. These methods control type I error rates, crucial for identifying complex disease risk variants.
Area of Science:
- Genetics
- Bioinformatics
- Statistical Genetics
Background:
- Binary phenotypes in biobank data often exhibit unbalanced case-control ratios, leading to inflated Type I error rates in association tests.
- Existing region-based tests struggle with extreme imbalance or scalability for large datasets.
Purpose of the Study:
- To develop accurate and scalable region-based tests for genetic association analysis with unbalanced binary phenotypes.
- To address limitations of current methods in controlling Type I error rates under extreme case-control imbalance.
Main Methods:
- Proposed SKAT- and SKAT-O- type region-based tests calibrating single-variant score statistics using saddle point approximation (SPA) and efficient resampling (ER).
- Simulation studies to evaluate Type I error rates and calibration of p-values.
- Application to UK Biobank whole-exome sequence data (45,596 samples, 791 phenotypes).
Main Results:
- The proposed method demonstrated well-calibrated p-values in simulations, contrasting with greatly inflated Type I error rates (90x exome-wide alpha) of unadjusted approaches at a 1:99 ratio.
- The method showed similar computational time to unadjusted approaches, confirming scalability for large-scale biobank data.
- Analysis of UK Biobank data identified 10 rare-variant associations (p < 10^-7), including JAK2/myeloproliferative disease, HOXB13/prostate cancer, and F11/congenital coagulation defects.
Conclusions:
- The novel SPA and ER-calibrated region-based tests provide accurate and scalable solutions for analyzing unbalanced biobank data.
- These methods enhance the identification of rare-variant associations with complex diseases.
- Publicly available results via a web server facilitate further genetic research.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Single Nucleotide Polymorphisms-SNPs

