Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

12.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
12.2K
Genomics02:02

Genomics

35.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
35.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Pilot Project Leveraging Large Language Models for Automated Screening and Variable Extraction in Observational Studies.

medRxiv : the preprint server for health sciences·2026
Same author

Detecting Uncoded Self-Harm in Veterans' Electronic Health Records Using Positive and Unlabeled Learning: Retrospective Cohort Study.

Journal of medical Internet research·2026
Same author

The Common Fund Data Ecosystem (CFDE).

bioRxiv : the preprint server for biology·2026
Same author

Effective spectrum-based antibiotic resistance index for monitoring resistance in Gram-negative bacilli.

Antimicrobial stewardship & healthcare epidemiology : ASHE·2026
Same author

KG2ML: integrating knowledge graphs and positive unlabeled learning for identifying disease-associated genes.

Frontiers in bioinformatics·2026
Same author

Badapple 2.0: An Empirical Predictor of Compound Promiscuity, Updated, Modernized, and Enhanced for Explainability.

Journal of chemical information and modeling·2025

Related Experiment Video

Updated: May 16, 2025

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

1.5K

KG2ML: Integrating Knowledge Graphs and Positive Unlabeled Learning for Identifying Disease-Associated Genes.

Praveen Kumar1, Vincent T Metzger1, Swastika T Purushotham1

  • 1University of New Mexico (UNM), School of Medicine, Department of Internal Medicine, Translational Informatics Division, Albuquerque, New Mexico, USA.

Medrxiv : the Preprint Server for Health Sciences
|April 1, 2025
PubMed
Summary

This study introduces KG2ML, a novel pipeline using Positive and Unlabeled (PU) learning to discover new disease-associated genes. KG2ML effectively identifies hidden gene-disease relationships, advancing biomedical research.

Keywords:
KG2MLPULSCARPULSNARPositive and Unlabeled (PU) learningProteinGraphMLgene-disease association

More Related Videos

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

614
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

10.1K

Related Experiment Videos

Last Updated: May 16, 2025

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

1.5K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

614
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

10.1K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Machine Learning

Background:

  • Biomedical knowledge graphs (KGs) like DDKG map known gene-disease relationships but miss unexplored associations.
  • Identifying novel disease-associated genes is crucial but challenging with traditional methods.
  • Machine learning (ML) and KGs offer promising avenues for inferring unknown associations.

Purpose of the Study:

  • To develop an efficient computational approach for predicting novel disease-associated genes.
  • To overcome limitations of existing ML pipelines for KG analysis.
  • To identify previously unknown gene-disease links.

Main Methods:

  • Developed KG2ML, a novel ML pipeline integrating Positive and Unlabeled (PU) learning (PULSNAR) with path-based feature extraction.
  • Applied KG2ML to 12 diseases to infer missing gene associations from the Data Distillery Knowledge Graph (DDKG).
  • Utilized ProteinGraphML for feature extraction and XGBoost for classification.

Main Results:

  • KG2ML identified potential disease-associated genes for 12 diseases, with many lacking prior explicit links in DDKG.
  • Top-ranked imputed genes were supported by literature and TINX evidence.
  • Incorporating PULSNAR-imputed genes improved XGBoost classification performance, validating PU learning's potential.

Conclusions:

  • PU learning, via KG2ML, effectively uncovers disease-gene associations absent in current KGs.
  • The KG2ML pipeline offers a scalable and interpretable framework for biomedical research.
  • This approach enhances KG utility and advances the discovery of novel gene-disease relationships.