莫卡:多omics桥接SNP集内核关联测试的管道
David Enoma1,2,3, Dinghao Wang4, Ariel Ghislain Kemogne Kamdoum4
1Department of Biochemistry and Molecular Biology, Cumming School of Medicine, University of Calgary, Calgary, AB T2N 4N1, Canada.
G3 (Bethesda, Md.)
|December 19, 2025
概括
我们开发了MOKA,这是一个集成多omics数据的新管道,用于全基因组关联研究 (GWAS). 这种工具增强了对精神分裂症等复杂疾病的变异发现和分析.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 越来越多的基因组和多基因组数据需要用于全基因组关联研究 (GWAS) 的先进工具.
- 整合功能注释对于增强GWAS的功能和可解释性至关重要.
- 当前的方法在处理各种功能数据类型时,往往缺乏可扩展性和可重复性.
研究的目的:
- 为了引入多omics数据桥接内核协会测试 (MOKA) 管道.
- 为GWAS提供可扩展和可重复的工作流程,将多omics数据集成到GWAS中.
- 改进遗传关联研究中的变异优先级和统计能力.
主要方法:
- 开发了MOKA,这是一个基于Snakemake的工作流程,用于基于SNP集内核的关联测试.
- 整合了多种多omics数据:基因表达,转录因子结合,保护得分和神经网络特征.
- 实施了人口结构校正,并行计算和全面的GWAS后分析 (可视化,GO注释,路径丰富).
主要成果:
- 将MOKA应用于精神分裂症GWAS队列,确定了89个邦费罗尼显著基因.
- 使用DisGeNET数据库实现了15.7%的验证率.
- 观察到与神经精神疾病相关的途径的丰富,证明了MOKA的实用性.
结论:
- MOKA提供了一个强大的,可扩展的,可扩展的框架,用于在遗传学研究中功能性多omics集成.
- 该管道增强了GWAS中的变量优先级和统计能力.
- MOKA是开源的,促进了在遗传研究中更广泛的采用.
相关概念视频
Single Nucleotide Polymorphisms-SNPs
17.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.8K
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Genome-wide Association Studies-GWAS
15.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.2K
Genomics
39.5K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
39.5K
DNA Microarrays
20.6K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
20.6K
Sanger Sequencing
772.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.8K


