Related Experiment Video
Updated: Jan 28, 2026

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
KG2ML: integrating knowledge graphs and positive unlabeled learning for identifying disease-associated genes
Praveen Kumar1, Vincent T Metzger1, Swastika T Purushotham1
1Department of Internal Medicine, Translational Informatics Division, School of Medicine, University of New Mexico (UNM), Albuquerque, NM, United States.
This study introduces KG2ML, a novel pipeline using Positive and Unlabeled (PU) learning to discover hidden disease-gene associations in biomedical knowledge graphs (KGs). The method successfully identified potential new disease-associated genes, enhancing research capabilities.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning
Background:
- Biomedical knowledge graphs (KGs) like DDKG store known relationships but miss unexplored associations.
- Identifying unknown disease-associated genes is critical for advancing biomedical research.
- Existing methods are time-consuming, necessitating efficient computational approaches.
Purpose of the Study:
- To develop an efficient computational approach for identifying novel disease-associated genes.
- To overcome limitations of existing machine learning pipelines for knowledge graph analysis.
- To infer previously unknown gene-disease relationships using advanced machine learning.
Main Methods:
- Developed KG2ML (Knowledge Graph to Machine Learning) pipeline, utilizing a novel Positive and Unlabeled (PU) learning algorithm, PULSCAR (Positive Unlabeled Learning Selected Completely At Random).
- Incorporated path-based feature extraction from ProteinGraphML within the KG2ML pipeline.
- Applied KG2ML to 12 diseases to infer novel disease-associated genes not present in the DDKG.
Main Results:
- KG2ML identified top-ranked novel disease-associated genes for 12 diseases, with 14 out of 15 lacking prior explicit associations in DDKG.
- Identified genes showed support from literature and TINX (Target Importance and Novelty Explorer) evidence.
- Incorporating PULSCAR-imputed genes as positives improved XGBoost classification performance.
Conclusions:
- Positive and Unlabeled (PU) learning effectively uncovers disease-gene associations missing from existing knowledge graphs (KGs).
- The KG2ML pipeline offers a scalable and interpretable framework for biomedical research.
- Integrating KG data with ML-based inference advances biomedical discovery by addressing KG limitations.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:58C. elegans Positive Butanone Learning, Short-term, and Long-term Associative Memory Assays
Published on: March 11, 2011
Related Concept Videos
Chromatin Position Affects Gene Expression
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
Velocity and Position by Integral Method
Consider an example to calculate the velocity and position from the acceleration function. A motorboat is traveling at a constant velocity of 5.0 m/s when it starts to decelerate to arrive at the dock. Its acceleration is...
Ogive Graph
Graphing Antiderivatives
Design Example: Identifying the Locations of Monuments in the Field Using Global Positioning System Device
Bar Graph