Related Experiment Video
Updated: Nov 25, 2025

07:35
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
1.9K
Graph embeddings on gene ontology annotations for protein-protein interaction prediction
Xiaoshi Zhong1, Jagath C Rajapakse2
1School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China. xszhong@ntu.edu.sg.
BMC Bioinformatics
|December 16, 2020
Summary
This study introduces a novel graph embedding method to address missing and spurious protein-protein interactions (PPIs). The approach effectively improves PPI prediction accuracy by learning from Gene Ontology Annotation graphs.
Area of Science:
- Bioinformatics
- Computational Biology
- Systems Biology
Background:
- Protein-protein interaction (PPI) prediction is crucial for understanding biological functions, gene-disease, and disease-drug associations.
- Existing PPI prediction methods often overlook missing and spurious interactions within PPI networks.
- Addressing these limitations is vital for accurate biological network analysis.
Purpose of the Study:
- To develop a method for predicting missing and spurious protein-protein interactions (PPIs).
- To leverage graph embeddings for learning vector representations from Gene Ontology Annotation (GOA) graphs.
- To enhance the accuracy of PPI network analysis by accounting for network imperfections.
Main Methods:
- Constructed Gene Ontology Annotation (GOA) graphs incorporating term-term relations and term-protein annotations.
- Employed graph embedding techniques to learn vector representations from the GOA graphs.
- Utilized these learned embeddings to perform missing and spurious PPI prediction tasks.
- Preserved both local and global structural information within the GOA graph.
Main Results:
- The proposed graph embedding method demonstrated superior performance compared to information content (IC)-based and word embedding-based methods.
- Experiments conducted on three PPI datasets from the STRING database validated the method's effectiveness.
- The approach successfully identified missing and spurious interactions in PPI networks.
Conclusions:
- Graph embeddings provide an effective means to learn vector representations from GOA graphs for PPI prediction.
- The method successfully addresses the challenges of missing and spurious interactions in PPI networks.
- This approach enhances the reliability and accuracy of bioinformatics predictions based on PPI data.
Related Concept Videos
Protein Networks
4.3K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.3K
Protein Networks
2.6K
2.6K
Protein-protein Interfaces
14.2K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.2K
Protein-Protein Interfaces
4.2K
4.2K
Genome Annotation and Assembly
19.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.9K
Protein Families
16.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.3K

