通过利用稀疏的祖先调整样本相关性来进行大型多祖先生物库的可扩展分析
Xihong Lin1, Rounak Dey2, Xihao Li3
1Harvard T.H. Chan School of Public Health.
Research square
|November 28, 2024
概括
FastSparseGRM改善了不同种群的遗传关联研究. 这种可扩展的管道精确地控制了种群结构和样本相关性,为全基因组关联研究 (GWAS) 和全基因组测序 (WGS) 提供了增强的功率和速度.
科学领域:
- 人口遗传学 人口遗传学
- 统计基因组学 统计基因组学
- 生物信息学是一种生物信息学.
背景情况:
- 线性混合效应模型 (LMMs) 和回归是控制遗传关联研究中人口结构和样本相关性的标准.
- 使用实证遗传亲属关系矩阵 (GRM) 的现有方法在多祖先种群中扎,导致混,膨胀的I型错误和减少功率.
研究的目的:
- 推出FastSparseGRM,这是一个可扩展的计算管道,用于多祖先的全基因组广泛关联研究 (GWAS) 和全基因组测序 (WGS).
- 通过提高准确性,速度和功率来解决异质群体中现有方法的局限性.
主要方法:
- 使用块对角的稀疏祖先调整 (BDSA) GRM来建模样本相关性.
- 纳入祖先主要组件 (PC) 作为固定效应来控制人口结构.
- 通过数值模拟和英国生物银行数据集中的五种生物标志物的分析来评估.
主要成果:
- 在大型异构队列中,FastSparseGRM在零LMM适配方面显示出显著的速度改进 (比BOLT-LMM/fast-GWA/REGENIE快约2540/4100/54倍) .
- 该方法有效地扩展到近50万名受试者.
- 与现有方法相比,实现精确的p值校准和增强的统计能力.
结论:
- FastSparseGRM提供了一个可扩展和准确的解决方案,用于多祖先种群的遗传关联研究.
- BDSA GRM和祖先PC有效地将人口结构从样本相关性中解脱出来,克服了传统GRM的局限性.
- 这种方法为GWAS和WGS提供了更好的功率和计算效率,促进了多样化的队列中的发现.
相关概念视频
Genome-wide Association Studies-GWAS
12.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.5K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Comparing Copy Number Variations and SNPs
17.3K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.3K
Pedigree Analysis
84.0K
Overview
84.0K
Gene Evolution - Fast or Slow?
7.0K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.0K
One-Way ANOVA: Unequal Sample Sizes
5.7K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.7K


