增强罕见变异关联研究的实力,通过使用大规模人口测序来进行推算
Jinglan Dai1, Yixin Zhang1, Yuan Gao1
1Department of Biostatistics, Center for Global Health, School of Public Health, Nanjing Medical University, Nanjing 211166, China.
Genomics, proteomics & bioinformatics
|September 19, 2025
概括
假定遗传数据,特别是TOPMed的数据,与罕见变异关联研究的全基因组测序 (WGS) 非常接近. 更大的样本大小和归算数据显著增加了检测罕见变异的能力,有助于复杂的特征遗传性研究.
科学领域:
- 基因组学就是基因组学.
- 人口遗传学 人口遗传学
- 统计遗传学 统计遗传学
背景情况:
- 种群规模的全基因组测序 (WGS) 能够精确捕获罕见变异,这对于理解复杂特征遗传性至关重要,通常被传统的全基因组关联研究 (GWAS) 遗漏.
- 假定遗传数据与WGS对于罕见变异关联研究的实用性仍然是一个开放的问题.
研究的目的:
- 评估SNP数组数据与WGS对比的统一性和性能,用于罕见变异检测.
- 评估各种特征和样本大小的关联研究中计入数据的力量.
主要方法:
- 利用英国生物银行提供的WGS数据 (n=150,119) 作为基础真相.
- 对于使用R平方和克莱默V的罕见变体,评估了归算质量 (TOPMed,HRC+UK10K).
- 对45个特征进行了关联测试,比较WGS和不同样本大小的归算数据.
- 使用SNP阵列和WGS数据对肺癌和卵巢癌进行了元分析.
主要成果:
- 即使对于极其罕见的变异 (小等位基数≤5) 实现了优质的TOPM归算 (R-平方>0.6).
- 跨种族的TOPMed输入数据显示了与WGS (克莱默V>0.75) 的更高一致性.
- 在相同的样本大小下,归算数据没有超过WGS,但TOPMed-归算数据结果更一致.
- 随着TOPMed输入数据的样本规模增加到488,377个,罕见变种检测率提高了27.71% (定量) 和10倍 (二进制).
- 分析发现了更多的变体和基因,而不是仅WGS的结果.
结论:
- 从大型测序队列中输入的罕见变异可以增强关联测试的功率,特别是当WGS样本大小有限时.
- TOPMed输入数据为罕见变异关联研究提供了宝贵的资源,接近WGS质量并通过规模提高发现能力.
相关概念视频
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K
RNA-seq
11.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.8K
Next-generation Sequencing
97.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.8K
Sanger Sequencing
773.3K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
773.3K


