iCircDA-NEAE:加速属性网络嵌入和动态卷积自编码器用于circRNA-疾病关联预测
Lin Yuan1,2,3, Jiawang Zhao1,2,3, Zhen Shen4
1Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), Jinan, China.
PLoS computational biology
|August 31, 2023
概括
这项研究介绍了iCircDA-NEAE,这是一种用于预测循环RNA (circRNA) -疾病关联的新型深度学习模型. 该模型有效地利用各种数据类型来提高预测准确性并识别潜在的疾病生物标志物.
科学领域:
- 生物医学信息学是生物医学信息学.
- 计算生物学是一种计算生物学.
- 基因组学就是基因组学.
背景情况:
- 循环RNAs (circRNAs) 越来越多地被认为是它们在人类疾病中的作用.
- 预测circRNA疾病关联有助于了解疾病机制,诊断和生物标志物发现.
- 现有的深度学习方法通常不充分利用生物识别数据,并提取次优特征.
研究的目的:
- 开发一种新的深度学习模型,iCircDA-NEAE,用于准确的circRNA疾病关联预测.
- 通过整合多种数据源和增强特征提取来解决当前方法的局限性.
- 为了改善临床应用的潜在circRNA疾病关系的识别.
主要方法:
- 开发了iCircDA-NEAE深度学习模型.
- 综合性疾病的语义相似性,高斯交互概况内核,circRNA表达概况相似性,和雅卡德相似性.
- 采用加速属性网络嵌入 (AANE) 和动态卷积自编码器 (DCAE) 来进行特征提取.
主要成果:
- 在circR2Disease数据集上,iCircDA-NEAE与现有方法相比,表现明显优越.
- 在前20个预测的circRNA-疾病对中,有16个通过现有文献得到了验证.
- 该模型有效地预测了新的潜在circRNA疾病关联.
结论:
- iCircDA-NEAE提供了一个强大的新工具,用于预测circRNA与疾病的关联.
- 该模型能够集成多种数据类型并提取强大的特征,从而增强其预测能力.
- 这种方法对推进疾病病原性研究和生物标志物发现有前途.
相关概念视频
RNA-seq
10.1K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.1K
Improving Translational Accuracy
11.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.6K
lncRNA - Long Non-coding RNAs
8.6K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.6K
Genome-wide Association Studies-GWAS
13.6K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.6K


