Related Experiment Video
Updated: Mar 21, 2026

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Generating Gene Ontology-Disease Inferences to Explore Mechanisms of Human Disease at the Comparative Toxicogenomics
Allan Peter Davis1, Thomas C Wiegers1, Benjamin L King2
1Department of Biological Sciences, North Carolina State University, Raleigh, North Carolina, United States of America.
Abstract:
Strategies for discovering common molecular events among disparate diseases hold promise for improving understanding of disease etiology and expanding treatment options. One technique is to leverage curated datasets found in the public domain. The Comparative Toxicogenomics Database (CTD; http://ctdbase.org/) manually curates chemical-gene, chemical-disease, and gene-disease interactions from the scientific literature. The use of official gene symbols in CTD interactions enables this information to be combined with the Gene Ontology (GO) file from NCBI Gene. By integrating these GO-gene annotations with CTD's gene-disease dataset, we produce 753,000 inferences between 15,700 GO terms and 4,200 diseases, providing opportunities to explore presumptive molecular underpinnings of diseases and identify biological similarities. Through a variety of applications, we demonstrate the utility of this novel resource. As a proof-of-concept, we first analyze known repositioned drugs (e.g., raloxifene and sildenafil) and see that their target diseases have a greater degree of similarity when comparing GO terms vs. genes. Next, a computational analysis predicts seemingly non-intuitive diseases (e.g., stomach ulcers and atherosclerosis) as being similar to bipolar disorder, and these are validated in the literature as reported co-diseases. Additionally, we leverage other CTD content to develop testable hypotheses about thalidomide-gene networks to treat seemingly disparate diseases. Finally, we illustrate how CTD tools can rank a series of drugs as potential candidates for repositioning against B-cell chronic lymphocytic leukemia and predict cisplatin and the small molecule inhibitor JQ1 as lead compounds. The CTD dataset is freely available for users to navigate pathologies within the context of extensive biological processes, molecular functions, and cellular components conferred by GO. This inference set should aid researchers, bioinformaticists, and pharmaceutical drug makers in finding commonalities in disease mechanisms, which in turn could help identify new therapeutics, new indications for existing pharmaceuticals, potential disease comorbidities, and alerts for side effects.
Insights
Researchers integrated public databases to uncover shared molecular events across diseases. This approach aids in understanding disease origins and developing new treatments by linking gene functions to various conditions.
Area of Science:
- Genomics
- Toxicology
- Bioinformatics
- Computational Biology
Background:
- Discovering common molecular events across diseases can advance etiological understanding and treatment strategies.
- Publicly available curated datasets offer a valuable resource for such discoveries.
Purpose of the Study:
- To integrate data from the Comparative Toxicogenomics Database (CTD) with Gene Ontology (GO) annotations to create a novel resource for exploring disease relationships.
- To demonstrate the utility of this resource in identifying disease similarities, predicting drug repositioning candidates, and generating therapeutic hypotheses.
Main Methods:
- Manually curated chemical-gene, chemical-disease, and gene-disease interactions from CTD were combined with NCBI Gene's GO-gene annotations.
- A large-scale inference set of GO terms and diseases was generated.
- Applications included analyzing drug repositioning, predicting disease similarities, and identifying potential drug candidates for specific cancers.
Main Results:
- Generated over 753,000 inferences linking 15,700 GO terms to 4,200 diseases.
- Demonstrated that diseases targeted by repositioned drugs show greater similarity when analyzed via GO terms compared to genes alone.
- Predicted and validated novel disease-disease similarities (e.g., stomach ulcers and atherosclerosis with bipolar disorder).
- Identified potential drug candidates, including cisplatin and JQ1, for B-cell chronic lymphocytic leukemia.
Conclusions:
- The integrated CTD and GO dataset provides a powerful resource for researchers to explore disease pathologies and molecular underpinnings.
- This resource facilitates the identification of common disease mechanisms, potential therapeutic targets, new drug indications, comorbidities, and potential side effects.
- The findings support the use of integrated biological data for advancing drug discovery and understanding complex diseases.
Related Concept Videos
Genomics
Pharmacogenomics: Identification of New Drug Targets
Pharmacogenetics and Pharmacogenomics: Overview
Mutagenicity and Carcinogenicity
Toxicity Testing in Animals
Principles of Pharmacogenetics: Types of Genetic Variants

