预测 lncRNA-疾病协会使用多个元路径在等级图注意力网络
Dengju Yao1, Yuexiao Deng2, Xiaojuan Zhan2,3
1School of Computer Science and Technology, Harbin University of Science and Technology, Harbin, 150080, China. ydkvictory@hrbust.edu.cn.
BMC bioinformatics
|January 29, 2024
概括
本研究介绍了MMHGAN,这是一种深度学习模型,通过分析复杂的网络结构,有效预测长非编码RNA (lncRNA) -疾病关联. 该模型的准确性很高,性能优于现有方法,有助于了解疾病的发病因子.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 长非编码RNAs (lncRNAs) 在调节基因表达中起着至关重要的作用,并与复杂疾病的发病有关.
- 预测lncRNA与疾病的关联对于了解疾病机制至关重要,但由于大量的lncRNA和实验的局限性,这具有挑战性.
- 现有的计算方法往往忽略了网络结构中中间节点提供的信息.
研究的目的:
- 开发一种新的深度学习模型,MMHGAN,用于预测未知的lncRNA-疾病关联.
- 为了利用层次化的图形注意网络和多个元路径类型来增强功能提取.
- 通过考虑更广泛的网络信息,包括中间节点来改进现有方法.
主要方法:
- 构建一个整合lncRNA-疾病-miRNA关联的异质图和lncRNA和疾病的同质图.
- 多头注意力机制的应用,用于在同质图中聚合特征.
- 在异质图中选择和权衡具有不同中间节点的元路,以获得最终的嵌入式特征.
主要成果:
- 通过五重交叉验证,通过五重交叉验证获得了平均AUC96.07%和平均AUPR93.23%.
- 废弃实验证实了同质图形和不同的中间节点路径权重的重要性.
- 肺癌,食道癌和乳腺癌的案例研究显示,预测的lncRNA与疾病相关性具有很高的验证率.
结论:
- 与六种现有的预测模型相比,MMHGAN模型显示出更高的性能.
- 该模型有效地预测了潜在的lncRNA-疾病相关性,并通过案例研究进行了验证.
- MMHGAN提供了一种有前途的计算方法,用于推进 lncRNA-疾病关联研究.
相关概念视频
lncRNA - Long Non-coding RNAs
8.6K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.6K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K


