Related Experiment Video
Updated: Jun 4, 2026

07:35
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Automated ontological gene annotation for computing disease similarity
Sachin Mathur1, Deendayal Dinakarpandian
1University of Missouri-Kansas City, Kansas City, Missouri.
Summit on Translational Bioinformatics
|February 25, 2011
Summary
This study enhances gene-disease associations by mapping protein data to the Disease Ontology (DO), improving clinical diagnosis and drug discovery. The new method reveals significantly more gene-disease links and disease similarities.
Area of Science:
- Genomics
- Bioinformatics
- Medical Informatics
Background:
- Accurate gene-disease association is crucial for clinical diagnosis and drug discovery.
- Existing methods primarily focus on gene expression or genetic data, with limited exploration of protein sequence data for disease terminology mapping.
- The Disease Ontology (DO) provides a valuable framework for organizing disease information.
Purpose of the Study:
- To augment the Disease Ontology (DO) with protein sequence data for automated gene-disease association.
- To improve the accuracy and scope of gene-disease annotations.
- To develop a method for measuring disease similarity based on protein co-occurrence and ontology structure.
Main Methods:
- Augmenting the Disease Ontology (DO) using protein sequence data from Swissprot records.
- Automated annotation of Swissprot records against the augmented DO.
- Measuring disease similarity by analyzing co-occurrence of annotations among proteins and leveraging the hierarchical structure of DO.
Main Results:
- Achieved a 36.1% increase in gene-disease associations compared to the original DO.
- Successfully mapped protein sequence data to disease terminologies.
- Enabled measurement of disease similarity, facilitating the identification of related diseases and potential novel relationships.
Conclusions:
- Augmenting the Disease Ontology with protein sequence data significantly enhances gene-disease associations.
- The developed method provides a robust approach for automated annotation and disease similarity measurement.
- This work has the potential to accelerate clinical diagnosis, drug discovery, and the identification of novel disease connections.
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Genomics
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...

