Related Experiment Video
Updated: Jan 9, 2026

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Prior knowledge informs graph neural networks to improve phenotype prediction from proteomics
Prabuddha Ghosh Dastidar1, Gus Fridell2, Joshua M Popp3
1Department of Computer Science, Johns Hopkins University, Baltimore, MD, USA.
Abstract:
High-throughput proteomics data provides dense individual-level molecular readouts, enabling the development of machine learning models for predicting diverse phenotypes relevant to patient health. Proteins interact in the cell in complex, nonlinear relationships that may not be reflected by linear models or simple machine learning approaches, highlighting the potential for more expressive deep neural networks to improve performance. Despite this possibility, in practice, developing neural network approaches in biological domains has been a significant challenge. We developed a deep learning framework for predicting disease-related traits from protein expression data using an innovative model architecture designed to exploit structured biological knowledge. The core of the model is a graph neural network (GNN) operating on bipartite graphs where one set of nodes represents protein expression levels and the other represents hundreds of protein sets derived from gene ontology libraries. Edges encode set membership, providing a compact and biologically meaningful structure. We trained our model using the UK Biobank plasma proteomics and individual phenotype data. Of the architectures we examined, the best-performing architecture had three parallel heads: two GNNs each using graphs constructed with independent protein set libraries and one global head consisting of tabular protein expression data. Their outputs are concatenated and passed through a dense feed-forward network to predict phenotype. When applied to predicting glycated hemoglobin (HbA1c) levels and a range of other phenotypes, our model showed strong predictive performance, outperforming other deep learning architectures and simpler linear models. Control models with permuted protein labels displayed worse performance demonstrating that the model benefits from the inductive bias from incorporating prior knowledge, especially in settings with limited training data. We present an innovative model architecture incorporating biological domain knowledge to predict complex traits from large scale proteomic data.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:28JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Protein-protein Interfaces
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Human Genetics
The complex relationship between genetics and psychology is observable through common biological components such...