通过交叉队列分析展示解锁云基因组学价值的途径
Nicole Deflaux1, Margaret Sunitha Selvaraj2,3,4,5, Henry Robert Condon6
1Verily Life Sciences, San Francisco, CA, USA.
Nature communications
|September 5, 2023
概括
大规模的基因组项目使用基于云的可信研究环境 (TREs). 对脂质特征进行比较的元分析和聚合分析揭示了不同的变体发现,特别是在不同的祖先中.
科学领域:
- 基因组学和生物信息学
- 人口遗传学 人口遗传学
- 数据科学在健康研究研究中的数据科学
背景情况:
- 基于云计算的可信研究环境 (TREs) 中的集中数据存储是大型基因组项目 (如All of Us和英国生物银行) 的新范式.
- 了解TRE属性对交叉队列分析的影响对于最大限度地提高数据实用性至关重要.
- 脂质测量是评估基因组分析方法的标准表型,因为之前进行了广泛的研究.
研究的目的:
- 在TREs中比较基因组广泛关联研究 (GWAS) 的元分析和聚合分析方法.
- 在交叉队列分析中描述TRE属性的优缺点.
- 评估分析选择对鉴定与脂质特征相关的遗传变异的影响,特别是在不同的祖先之间.
主要方法:
- 进行了一项全基因组关联研究 (GWAS),用于标准脂质测量,使用元分析和聚合分析.
- 利用两种基于TRE的方法的总结数据,并将结果与外部研究进行了比较.
- 评估了已识别的基因位点与脂质水平的相关性,并评估了祖先群体的变异意义和流行率.
主要成果:
- 无论是元分析还是聚合分析,都显示出与已知的脂质位置 (R2 ≈ 83-97%) 有着强烈的相关性.
- 有相当数量的变异在元分析 (90种变异) 或聚合分析 (64种变异) 中完全达到全基因组显著性值.
- 大约20%的这些独特的显著变异在非欧洲,非亚洲祖先的个体中最为普遍,这表明了祖先特定的发现.
结论:
- 在TREs中的技术和政策决策可能会导致跨队列分析的不同结果,即使使用相似的数据.
- 分析方法的差异 (元分析与聚合分析) 影响了遗传变异的检测,特别是代表性不足的祖先群体.
- 未来的交叉队列分析必须仔细考虑分析方法,以确保在不同人群中进行公平的发现,避免加剧健康差异.
相关概念视频
Genomics
36.4K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.4K
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Genome-wide Association Studies-GWAS
13.6K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.6K
DNA Microarrays
17.5K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.5K
Next-generation Sequencing
91.4K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
91.4K


