在GWAS中纠正志愿者偏差会增加SNP效应大小和遗传概率估计
Sjoerd van Alten1,2, Benjamin W Domingue3, Jessica Faul4
1Vrije Universiteit Amsterdam, Amsterdam, Netherlands. s.j.d.van.alten@vu.nl.
Nature communications
|April 15, 2025
概括
志愿者偏见显著影响全基因组关联研究 (GWAS). 这项研究引入了反向概率加权GWAS (WGWAS) 来纠正像英国生物银行这样的大队伍中的志愿者偏差,揭示了改变的遗传洞察力.
科学领域:
- 遗传学 是一个遗传学.
- 人口遗传学 人口遗传学
- 生物信息学是一种生物信息学.
背景情况:
- 在大型遗传研究中,如英国生物库 (UKB) 中,以志愿者为基础的抽样引入了选择偏差,通常称为志愿者偏差.
- 志愿者偏见对全基因组关联研究 (GWAS) 的程度和影响仍然不完全理解,可能会影响遗传发现.
研究的目的:
- 开发和应用一种反向概率加权GWAS (WGWAS) 方法,以纠正英国生物银行志愿者偏差的GWAS总结统计数据.
- 评估志愿者偏见纠正对遗传架构估计的影响,并确定受影响的表型.
主要方法:
- 使用英国人口普查数据估计的逆概率 (IP) 权重,确保英国生物银行采样人口的代表性.
- 应用IP权重,在英国生物银行内对10种表型进行WGWAS.
- 评估了IP权重的基于SNP的遗传性,以确认其捕获志愿者偏差的能力.
主要成果:
- 知识产权权重表现出显著的SNP-based遗传性 (4.8%),表明它们有效地捕捉了志愿者偏见.
- 与标准GWAS相比,WGWAS在10种表型中产生了更大的SNP效应大小和遗传性估计,尽管有效样本大小平均下降了62%.
- 志愿者偏见的影响因表型而异,疾病,健康行为和社会经济地位特征受到影响最大.
结论:
- 志愿者偏见显著扭曲了GWAS结果,特别是在某些特征类别.
- WGWAS提供了一种方法来调整GWAS总结统计数据以适应志愿者偏见,从而获得更准确的遗传见解.
- 建议GWAS联盟要么为其数据集提供人口权重,要么使用人口代表性样本来减轻志愿者偏见.
相关概念视频
Genome-wide Association Studies-GWAS
12.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.1K
Comparing Copy Number Variations and SNPs
16.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
16.8K
Single Nucleotide Polymorphisms-SNPs
13.6K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
13.6K
One-Way ANOVA: Equal Sample Sizes
3.1K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.1K
Hardy-Weinberg Principle
71.2K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
71.2K
Bias
3.7K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.7K


