NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association Prediction

Insights

This study introduces NSECDA, a novel method for predicting circRNA-disease associations (CDA). By treating circRNA sequences as biological language, it enhances disease diagnosis and pathogenesis insights.

Area of Science:

  • Biomarker Discovery
  • Genomics
  • Computational Biology

Background:

  • Circular RNAs (circRNAs) are emerging biomarkers with significant disease associations.
  • Existing circRNA-disease association (CDA) prediction models often overlook attribute influence and inter-attribute correlations.
  • There is a need for advanced models to leverage latent semantic information in circRNA and disease data.

Purpose of the Study:

  • To propose a natural semantic enhancement method (NSECDA) for predicting circRNA-disease associations (CDA).
  • To improve the accuracy and interpretability of CDA prediction by incorporating natural language understanding principles.
  • To identify novel CDA pairs for disease diagnosis and understanding pathogenesis.

Main Methods:

  • CircRNA sequences were analyzed using natural language understanding (NLU) theory to extract natural semantic properties.
  • Graph Attention Network (GAT) was employed to focus on influential attributes by integrating circRNA, disease attributes, and Gaussian Interaction Profile (GIP) kernel attributes.
  • Rotation Forest (RoF) classifier was utilized for the final prediction of CDA.

Main Results:

  • The NSECDA model achieved 92.49% accuracy and a 0.9225 AUC score on the CircR2Disease dataset.
  • NSECDA demonstrated competitive performance compared to non-enhanced models and other classifiers.
  • 25 of the top 30 predicted CDA pairs, previously unknown, were validated by recent studies.

Conclusions:

  • NSECDA is an effective model for predicting circRNA-disease associations.
  • The method provides credible candidates for wet experiments, significantly narrowing investigation scope.
  • This approach offers novel perspectives for disease diagnosis and understanding pathogenesis.

Related Concept Videos

Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
14.1K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.7K
RNA-seq03:21

RNA-seq

RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.3K