多元财政框架-nDA:基于多模块融合的ncRNA-疾病关联预测的计算模型
Zhihao Guan1,2, Xiu Jin1,2, Xiaodan Zhang1,2
1College of Information and Artificial Intelligence, Anhui Agricultural University, Hefei 230036, China.
Journal of chemical information and modeling
|March 25, 2025
概括
本研究介绍了MFF-nDA,这是一种用于预测非编码RNA与疾病相关性的新型计算模型. 它有效地整合了各种数据源,优于现有的方法,可以改善疾病机制的洞察力.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 非编码RNA (ncRNAs) 在基因调节和疾病中起着至关重要的作用.
- 准确预测ncRNA与疾病的关联对于理解疾病机制和开发疗法至关重要.
- 目前的计算模型,通常是基于GNN的,在整合全面的ncRNA和疾病信息方面存在局限性,并且可能遭受过度平滑.
研究的目的:
- 开发一个先进的计算模型,MFF-nDA,用于准确预测ncRNA-疾病关联.
- 通过整合多源特征信息和采用多模块融合方法来克服现有模型的局限性.
- 为了提高各种类型的ncRNAs的预测模型的概括性.
主要方法:
- 多年财政框架-nDA模型使用五种类型的相似性网络信息 (三种为ncRNA,两种为疾病).
- 它包含三个不同的模块:基于变压器的异质网络表示模块,基于GCN的关联网络模块和基于GAT的拓结构模块.
- 多模块融合学习用于捕获各种实体特征和拓信息.
主要成果:
- 在所有测试的数据集中,MFF-nDA模型实现了曲线下的面积 (AUC) 大于0.9000.
- 该模型与现有的最先进的方法相比,显示出更高的预测性能.
- 三个模块的互补效应有助于缓解某些GNN模型固有的过度平滑问题.
结论:
- 多年财政框架-nDA提供了一个非常准确和有价值的工具,用于识别潜在的ncRNA-疾病关联.
- 多模块融合方法有效地捕捉了复杂的关系,推进了ncRNA-疾病关联预测领域.
- 这种模型对阐明疾病机制和开发治疗策略有重大影响.
相关概念视频
RNA-seq
9.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.7K
Tagging and Fusion Proteins
6.6K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.6K
CRISPR and crRNAs
16.4K
Bacteria and archaea are susceptible to viral infections just like eukaryotes; therefore, they have developed a unique adaptive immune system to protect themselves. Clustered regularly interspaced short palindromic repeats and CRISPR-associated proteins (CRISPR-Cas) are present in more than 45% of known bacteria and 90% of known archaea.
The CRISPR-Cas system stores a copy of foreign DNA in the host genome and uses it to identify the foreign DNA upon reinfection. CRISPR-Cas has three different...
The CRISPR-Cas system stores a copy of foreign DNA in the host genome and uses it to identify the foreign DNA upon reinfection. CRISPR-Cas has three different...
16.4K
lncRNA - Long Non-coding RNAs
8.4K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.4K
Mouse Models of Cancer Study
5.5K
Mice have long served as models for studying human biology and pathology because of their phylogenetic and physiological similarity with humans. They are also easy to maintain and breed in the laboratory, and hence, many inbred strains are now available for research. Studies on mice have contributed immeasurably to our understanding of cancer biology.
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
5.5K
Single Nucleotide Polymorphisms-SNPs
13.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
13.8K


