改进的多祖先精细映射识别了底层分子特征和疾病风险的cis调控变异
Zeyun Lu1, Xinran Wang1, Matthew Carr1
1Center for Genetic Epidemiology, Keck School of Medicine, University of Southern California, Los Angeles, CA, USA.
概括
新的SuShiE模型提高了在多种祖先中识别因果性cis-molecular定量特征位点 (cis-molQTLs) 的精度. 它通过更好地反映共享的遗传架构来改善复杂特征的基因发现.
科学领域:
- 遗传学 是一个遗传学.
- 统计遗传学 统计遗传学
- 生物信息学是一种生物信息学.
背景情况:
- 对cis-molecular定量特征位点 (cis-molQTLs) 的统计精细映射旨在准确识别因果变异.
- 现有的方法无法充分考虑祖先之间共享的遗传架构.
- 杆链接不平衡 (LD) 异质性对于提高精细映射精度至关重要.
研究的目的:
- 引入共享单一效应总和 (SuShiE) 模型,用于多祖先的cis-molQTL精细映射.
- 为了提高精细映射精度,推断跨祖先效应大小相关性,并估计祖先特定的表达式预测重量.
- 识别与分子特征和复杂疾病风险相关的遗传变异.
主要方法:
- 将SuShiE模型应用于来自不同祖先的mRNA表达和血蛋白水平 (TOPMed MESA和GENOA研究).
- 与基线方法相比,SuShiE性能在基因发现和变异优先级方面进行了比较.
- 利用估计的cis-molQTL效应大小来进行全转录组 (TWAS) 和全蛋白质组 (PWAS) 关联研究.
主要成果:
- 与基线相比,SuShiE精制了16%更多的cis-molQTL,并优先考虑了较少的变体,功能丰富度更高.
- 平均而言,在祖先之间推断出一致的cis-molQTL架构,在预测功能丧失不容忍的基因中观察到显著的异质性.
- 使用SuShiE的TWAS和PWAS在一个大型生物银行队列中确定了44个与白细胞特征相关的额外基因.
结论:
- SuShiE显著提高了多祖先 cis-molQTL 精细映射的精度和功率.
- 该模型为跨祖先分子特征的共享和异质遗传架构提供了新的见解.
- SuShiE增强了与复杂疾病风险相关的基因的识别,证明了其在遗传关联研究中的实用性.
相关概念视频
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K
Cis-regulatory Sequences
9.9K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
9.9K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Single Nucleotide Polymorphisms-SNPs
15.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.0K
Incomplete Dominance
22.5K
Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
22.5K


