Related Experiment Video
Updated: May 15, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Interpretable candidate drug prioritization and explanation framework across-medical knowledge graphs based on graph
1Institute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, Beijing, China.
Objective:
Addressing the challenges in elucidating the mechanisms of complex diseases such as Type 2 Diabetes Mellitus (T2DM), this study aims to construct a domain-specific cross-medicine knowledge graph (CMKG) and develop a unified path scoring framework that couples graph embeddings with rule-based reasoning, enabling high-precision, interpretable prioritization and explanation of potential drug candidates.
Methods:
First, multi-source biomedical data from Hetionet, SymMap, TCMBank, STRING, and TTD were integrated. Using Jaccard and overlap-based fusion strategies, entity alignment and relation consolidation were performed to construct a deep CMKG bridged by genes. Second, four graph embedding models (TransE, DistMult, ComplEx, and RotatE) were introduced for link prediction and evaluated using MRR and Hits@K. Finally, to overcome the interpretability limitations of black-box predictions, AnyBURL rule learning was combined with depth-first search (DFS). We innovatively introduced an Ingredient Specificity Index (ISI) and a hybrid path confidence calibration mechanism, constructing a unified path scoring system incorporating length decay, node/relation weights, and experimental evidence bonuses to screen the most critical mechanistic paths.
Results:
The constructed CMKG contains 15 entity types (245,235 entities) and 52 relation types (7,155,373 triples), covering 709 core T2DM genes. Link prediction stability tests across multiple random seeds showed that the ComplEx model consistently performed best in handling complex multi-mapping relations (MRR = 0.213 ± 0.004, Hits@10 = 0.418 ± 0.003). Consequently, the fully converged ComplEx model (Peak Hits@10 = 0.48) was utilized for comprehensive prediction. Retaining the top 100 predictions, Abelmoschus manihot and Topiramate ranked highest among TCM herbs and modern medicine compounds, respectively. Path analysis based on the scoring system revealed deep multi-target mechanisms, including insulin signaling sensitization, inflammatory regulation, and chromatin/cell-cycle intervention.
Conclusion:
The proposed gene-bridged graph embedding and unified path scoring framework successfully translates probabilistic predictions into biologically traceable semantic explanations. Rigorous ablation and parameter sensitivity experiments confirm that the framework achieves a robust balance between knowledge coverage and explanatory specificity, providing a transparent, robust, and scalable methodological foundation for candidate drug prioritization in complex diseases.
Related Concept Videos
Pharmacogenomics: Identification of New Drug Targets
Diabetes Mellitus: Type 2 and Gestational
Diabetes Mellitus: Overview and Type I Subtype
Type 1 diabetes is an autoimmune disease in which the immune system mistakenly attacks and destroys the insulin-producing beta cells in the pancreas. As a result, the body is unable to produce sufficient insulin, and individuals with...
Fundamental Mathematical Principles in Pharmacokinetics: Calculus and Graphs
On the other hand, integral calculus focuses on...
Pharmacogenetics of Drug Metabolism: Overview
