跨多个组织的多维拼接事件的张量分解,以识别与复杂特征相关的拼接介导风险基因
Yan Yan1, Rui Chen2,3, Hakmook Kang1
1Department of Biostatistics, Vanderbilt University, Nashville, Tennessee, United States of America.
PLoS computational biology
|July 21, 2025
概括
这项研究引入了多组织剪接基因 (MTSG) 模型,以识别阿尔茨海默病和精神分裂症等复杂特征的剪接介导风险基因. MTSG有效地分析了多种组织的拼接变异,发现了与这些疾病的新型遗传联系.
科学领域:
- 遗传学 遗传学 是一个
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 识别复杂的特征风险基因是一项挑战.
- 转录组广泛关联研究 (TWAS) 方法整合基因表达和GWAS数据,以找到候选风险基因.
- 拼接,一个重要的遗传性贡献者,由于其复杂性,未被充分探索.
研究的目的:
- 开发一种新的方法,多组织拼接基因 (MTSG),用于分析跨多种组织的高维拼接事件.
- 通过将MTSG模型应用于GWAS数据来识别阿尔茨海默病 (AD) 和精神分裂症 (SCZ) 的剪接中介风险基因.
主要方法:
- 使用张量分解和稀疏的正规相关性分析 (sCCA) 开发了MTSG模型.
- 使用GTEx数据构建MTSG模型.
- 将MTSG模型应用于AD和SCZ的GWAS总结统计.
主要成果:
- 确定了174个重要的拼接介导的AD风险基因和497个SCZ.
- MTSG模型揭示了AD相关的途径和独特的AD风险基因,这些基因在单个组织分析中被遗漏了.
- 对SCZ的全脑MTSG模型显示,SCZ相关基因的丰富性更强,并确定了独特的风险基因.
结论:
- MTSG模型有效地捕捉了多组织拼接变异,提供了单组织方法所错过的见解.
- 开发的MTSG框架可用于发现各种复杂特征的剪接介导风险基因.
- 这种方法提高了复杂疾病的遗传基础的识别.
相关概念视频
Gene Duplication and Divergence
6.3K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.3K
Genome-wide Association Studies-GWAS
14.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.3K
Single Nucleotide Polymorphisms-SNPs
15.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.9K
Pleiotropy
41.1K
Pleiotropy is the phenomenon in which a single gene impacts multiple, seemingly unrelated phenotypic traits. For example, defects in the SOX10 gene cause Waardenburg Syndrome Type 4, or WS4, which can cause defects in pigmentation, hearing impairments, and an absence of intestinal contractions necessary for elimination. This diversity of phenotypes results from the expression pattern of SOX10 in early embryonic and fetal development. SOX10 is found in neural crest cells that form melanocytes,...
41.1K
Epistasis Analysis
5.2K
Although Mendel chose seven unrelated traits in peas to study gene segregation, most traits involve multiple gene interactions that create a spectrum of phenotypes. When the interaction of various genes or alleles at different locations influences a phenotype, this is called epistasis. Epistasis often involves one gene masking or interfering with the expression of another (antagonistic epistasis). Epistasis often occurs when different genes are part of the same biochemical pathway. The...
5.2K
RNA Splicing
57.1K
Splicing is the process by which eukaryotic RNA is edited before its translation into protein. The RNA strand transcribed from eukaryotic DNA is called the primary transcript. The primary transcripts that become mRNAs are called precursor messenger RNAs (pre-mRNAs). Eukaryotic pre-mRNA contains alternating sequences of exons and introns. Exons are nucleotide sequences that code for proteins, whereas introns are the non-coding regions. In RNA splicing, introns are removed and exons are bonded...
57.1K


