Related Experiment Video
Updated: Sep 3, 2025

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Implications of topological imbalance for representation learning on biomedical knowledge graphs
Stephen Bonner1, Ufuk Kirik2, Ola Engkvist3
1Data Sciences and Quantitative Biology, Discovery Sciences, R&D, AstraZeneca, Cambridge, UK.
Machine learning-driven drug discovery knowledge graphs (KGs) can be biased by data and modeling, causing overrepresented genes to rank highly regardless of biological relevance. Careful KG construction and interpretation are crucial for accurate gene-disease association predictions.
Area of Science:
- Computational biology
- Machine learning in drug discovery
- Bioinformatics
Background:
- Drug discovery knowledge graphs (KGs) leverage interconnected data for tasks like predicting gene-disease associations.
- Graph embedding methods (KGE) offer intuitive representations and inference capabilities.
- Ensuring biological meaningfulness of KG predictions is critical.
Purpose of the Study:
- To investigate biases in KG construction and their impact on gene ranking for disease association.
- To demonstrate how topological overrepresentation affects entity ranking in KGE models.
- To highlight the influence of entity frequency versus biological information in KGE predictions.
Main Methods:
- Utilized various datasets, KGE models, and predictive tasks to assess graph imbalances.
- Performed graph perturbation experiments to analyze model behavior.
- Evaluated the impact of structural imbalances on the ranking of genes for diseases.
Main Results:
- Demonstrated that densely connected entities are consistently ranked highly, irrespective of context.
- Provided evidence across diverse datasets and models that KGEs can prioritize entity frequency over biological relationships.
- Graph perturbation experiments confirmed the susceptibility of KGE models to structural biases.
Conclusions:
- Inherent structural imbalances in KGs can lead to biased predictions, overranking hubs.
- KGE models may rely more on entity frequency than biological relevance, impacting drug discovery.
- Practitioners must be aware of these modeling choices and biases when composing KGs and interpreting results.
More Related Videos
Related Concept Videos
Biostatistics: Overview
Discrete variables are...
Data: Types and Distribution
Distributions in...
Imbalances in Cardiac Output
CHF can occur due to the failure of either side of the heart. Left-side failure leads to pulmonary congestion—the right side continues to send...
Homeostatic Imbalance
However, sometimes these feedback loops fail,...
Bioequivalence: Overview
Improving Translational Accuracy

