Association studies with imputed variants using expectation-maximization likelihood-ratio tests.
Kuan-Chieh Huang1, Wei Sun2, Ying Wu3
1Department of Biostatistics, University of North Carolina, Chapel Hill, North Carolina, United States of America.
Plos One
|November 11, 2014
Summary
New methods improve genotype imputation uncertainty analysis in genetic studies. These approaches enhance statistical power and efficiency for association studies, especially with low-frequency variants.
Area of Science:
- Genetics
- Statistical Genetics
- Bioinformatics
Background:
- Genotype imputation is crucial in genetic studies, but low-quality imputation of low-frequency variants poses challenges.
- Growing reference panels improve imputation for many markers, yet increase the number of poorly imputed ones.
Purpose of the Study:
- To develop novel methods for incorporating genotype imputation uncertainty into association analyses.
- To enhance statistical power and computational efficiency in genetic association studies.
Main Methods:
- Developed an expectation-maximization likelihood-ratio test (EM-LRT) for association using posterior genotype probabilities (Scenario I).
- For imputed dosages (Scenario II), sampled genotype probabilities from their posterior distribution and applied EM-LRT.
- Validated methods through simulations and applications to real genetic datasets.
Main Results:
- Proposed EM-LRT methods maintain Type I error rates under both scenarios.
- EM-LRT-Prob (Scenario I) demonstrates optimal statistical power across various minor allele frequencies (MAF) and imputation qualities.
- EM-LRT-Dose (Scenario II) matches EM-LRT-Prob's power and surpasses standard methods for low MAF/imputation quality markers.
Conclusions:
- The proposed EM-LRT methods effectively handle imputation uncertainty in genetic association studies.
- These methods offer improved power and efficiency, particularly for challenging markers.
- Validates the utility of incorporating imputation uncertainty for robust genetic discoveries.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
17.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
17.3K
Probability Laws
45.5K
Overview
45.5K
Comparing Copy Number Variations and SNPs
19.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.4K
Multiple Allele Traits
39.3K
The Concept of Multiple Allelism
39.3K
Expected Frequencies in Goodness-of-Fit Tests
8.9K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.9K
Mechanistic Models: Compartment Models in Individual and Population Analysis
345
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
345


