在案例控制全基因组关联研究中的数学界限和效果大小
Sanjana M Paye1, Michael D Edge1
1Department of Quantitative and Computational Biology, University of Southern California.
bioRxiv : the preprint server for biology
|January 7, 2025
概括
在全基因组关联研究 (GWAS) 中优化病例分数对于统计能力至关重要. 我们的模型显示了不同病例数量如何影响检测遗传关联,这取决于等位基因特征.
科学领域:
- 人口遗传学 人口遗传学
- 统计遗传学 统计遗传学
- 基因组流行病学 基因组流行病学
背景情况:
- 病例控制全基因组关联研究 (GWAS) 对于识别与疾病相关的遗传变异至关重要.
- 研究设计决策,特别是病例与对照的比例,显著影响了GWAS的统计能力.
- 同基因频率和链接不平衡 (LD) 影响关联统计,而这些都受到病例分数的影响.
研究的目的:
- 调查在病例控制中的病例比例的变化如何影响GWAS检测遗传关联的统计能力.
- 扩展现有的知识,以了解LD统计的边界,以了解病例分数对功率的影响.
- 为优化基于等位基特征的研究设计提供框架.
主要方法:
- 一个数学模型的分析,其中包含了等位基因频率和LD.
- 模拟以评估与关联测试的非中心性参数成比例的数量.
- 探讨在不同条件下的影响主导地位,透率和等位基因频率.
主要成果:
- 案例分数显著影响非中心性参数,从而影响统计能力.
- 病例分数对功率的影响取决于风险等位基因的特定遗传结构 (主导性,透性,频率).
- 对风险与保护性等位基相比观察到的功率不对称性以及对某些等位基类型的平衡样本的非最佳功率得到解释.
结论:
- 在GWAS中最佳病例分数并不总是1:1并且取决于变异的遗传特性.
- 该框架提供了对优化GWAS设计的见解,以加强与疾病相关的遗传变异的检测.
- 这些发现适用于各种关联测试中的统计能力指南,超出了基平方测试.
更多相关视频
相关概念视频
Genome-wide Association Studies-GWAS
12.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.4K
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
119
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
119
Hardy-Weinberg Principle
71.7K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
71.7K
Accuracy and Errors in Hypothesis Testing
174
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
174
Chi-square Analysis
37.3K
The chi-square test is a statistical hypothesis test. It is used to check whether there is a significant difference between an expected value and an observed value. In the context of genetics, it enables us to either accept or reject a hypothesis, based on how much the observed values deviate from the expected values.
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
37.3K
Comparing Copy Number Variations and SNPs
17.1K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.1K


