Related Experiment Video
Updated: Jul 12, 2025

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Alexander R Pelletier1, Dylan Steinecke2, Dibakar Sigdel3
1Department of Physiology, UCLA School of Medicine; Scalable Analytics Institute (ScAi) at Department of Computer Science, UCLA School of Engineering; NIH BRIDGE2AI Center at UCLA & NHLBI Integrated Cardiovascular Data Science Training Program, UCLA; arpelletier@g.ucla.edu.
Abstract:
The rapidly increasing and vast quantities of biomedical reports, each containing numerous entities and rich information, represent a rich resource for biomedical text-mining applications. These tools enable investigators to integrate, conceptualize, and translate these discoveries to uncover new insights into disease pathology and therapeutics. In this protocol, we present CaseOLAP LIFT, a new computational pipeline to investigate cellular components and their disease associations by extracting user-selected information from text datasets (e.g., biomedical literature). The software identifies sub-cellular proteins and their functional partners within disease-relevant documents. Additional disease-relevant documents are identified via the software's label imputation method. To contextualize the resulting protein-disease associations and to integrate information from multiple relevant biomedical resources, a knowledge graph is automatically constructed for further analyses. We present one use case with a corpus of ~34 million text documents downloaded online to provide an example of elucidating the role of mitochondrial proteins in distinct cardiovascular disease phenotypes using this method. Furthermore, a deep learning model was applied to the resulting knowledge graph to predict previously unreported relationships between proteins and disease, resulting in 1,583 associations with predicted probabilities >0.90 and with an area under the receiver operating characteristic curve (AUROC) of 0.91 on the test set. This software features a highly customizable and automated workflow, with a broad scope of raw data available for analysis; therefore, using this method, protein-disease associations can be identified with enhanced reliability within a text corpus.
Related Concept Videos
Genomics
Lysosomal Hydrolases
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Smooth Endoplasmic Reticulum
The ER provides optimal conditions for synthesizing steroid hormones and lipids, such as phospholipids and triglycerides. Traditionally, lipid metabolism was considered to be a smooth ER function. However, there is no direct evidence to prove that rough ER is completely excluded from lipid...
Role of ER in the Secretory Pathway
Components of the secretory pathway
About a third of proteins synthesized in the cell are sorted via the secretory route. They shuffle between different compartments in membrane-bound vesicles until they reach their final destination. The main intracellular compartments involved...
Protein Import into the Peroxisomes
Peroxisomal Protein Import:
Peroxisomes lack the genetic machinery required to code for their own proteins. Hence, most peroxisomal membrane, lumenal and transmembrane proteins are synthesized in the cytoplasm or ER and transported to the peroxisome...

