Related Experiment Video
Updated: Sep 1, 2025

In Silico Identification and Characterization of circRNAs During Host-Pathogen Interactions
Published on: October 21, 2022
NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association Prediction
Insights
This study introduces NSECDA, a novel method for predicting circRNA-disease associations (CDA). By treating circRNA sequences as biological language, it enhances disease diagnosis and pathogenesis insights.
Area of Science:
- Biomarker Discovery
- Genomics
- Computational Biology
Background:
- Circular RNAs (circRNAs) are emerging biomarkers with significant disease associations.
- Existing circRNA-disease association (CDA) prediction models often overlook attribute influence and inter-attribute correlations.
- There is a need for advanced models to leverage latent semantic information in circRNA and disease data.
Purpose of the Study:
- To propose a natural semantic enhancement method (NSECDA) for predicting circRNA-disease associations (CDA).
- To improve the accuracy and interpretability of CDA prediction by incorporating natural language understanding principles.
- To identify novel CDA pairs for disease diagnosis and understanding pathogenesis.
Main Methods:
- CircRNA sequences were analyzed using natural language understanding (NLU) theory to extract natural semantic properties.
- Graph Attention Network (GAT) was employed to focus on influential attributes by integrating circRNA, disease attributes, and Gaussian Interaction Profile (GIP) kernel attributes.
- Rotation Forest (RoF) classifier was utilized for the final prediction of CDA.
Main Results:
- The NSECDA model achieved 92.49% accuracy and a 0.9225 AUC score on the CircR2Disease dataset.
- NSECDA demonstrated competitive performance compared to non-enhanced models and other classifiers.
- 25 of the top 30 predicted CDA pairs, previously unknown, were validated by recent studies.
Conclusions:
- NSECDA is an effective model for predicting circRNA-disease associations.
- The method provides credible candidates for wet experiments, significantly narrowing investigation scope.
- This approach offers novel perspectives for disease diagnosis and understanding pathogenesis.
Abstract:
Increasing evidence suggest that circRNA, as one of the most promising emerging biomarkers, has a very close relationship with diseases. Exploring the relationship between circRNA and diseases can provide novel perspective for diseases diagnosis and pathogenesis. The existing circRNA-disease association (CDA) prediction models, however, generally treat the data attributes equally, do not pay special attention to the attributes with more significant influence, and do not make full use of the correlation and symbiosis between attributes to dig into the latent semantic information of the data. Therefore, in response to the above problems, this paper proposes a natural semantic enhancement method NSECDA to predict CDA. In practical terms, we first recognize the circRNA sequence as a biological language, and analyze its natural semantic properties through the natural language understanding theory; then integrate it with disease attributes, circRNA and disease Gaussian Interaction Profile (GIP) kernel attributes, and use Graph Attention Network (GAT) to focus on the influential attributes, so as to mine the deeply hidden features; finally, the Rotation Forest (RoF) classifier was used to accurately determine CDA. In the gold standard data set CircR2Disease, NSECDA achieved 92.49% accuracy with 0.9225 AUC score. In comparison with the non-natural semantic enhancement model and other classifier models, NSECDA also shows competitive performance. Additionally, 25 of the CDA pairs with unknown associations in the top 30 prediction scores of NSECDA have been proven by newly reported studies. These achievements suggest that NSECDA is an effective model to predict CDA, which can provide credible candidate for subsequent wet experiments, thus significantly reducing the scope of investigations.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Single Nucleotide Polymorphisms-SNPs
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...

