Related Experiment Video
Updated: Jul 13, 2026

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
False positive reduction in protein-protein interaction predictions using gene ontology annotations
Mahmoud A Mahdavi1, Yen-Han Lin
1Department of Chemical Engineering, University of Saskatchewan, Saskatoon, SK, Canada. maa943@mail.usask.ca <maa943@mail.usask.ca>
BMC Bioinformatics
|July 25, 2007
Summary
This study reduces false positive protein-protein interactions (PPI) in computational predictions using Gene Ontology annotations and knowledge rules. This improves the accuracy and reliability of predicted PPI datasets for biological research.
Area of Science:
- Computational Biology
- Bioinformatics
- Systems Biology
Background:
- Protein-protein interactions (PPIs) are fundamental to cellular processes like metabolism and signaling.
- A significant challenge in PPI research is the low agreement between experimental and computational predictions due to high false positive rates in computational methods.
- Improving the accuracy of predicted PPIs by reducing false positives using reliable experimental data remains an underexplored area.
Purpose of the Study:
- To reduce false positive protein-protein interaction (PPI) pairs generated by computational prediction methods.
- To enhance the true positive fraction and robustness of computationally predicted PPI datasets.
- To leverage Gene Ontology (GO) annotations and experimental PPI data to improve prediction accuracy.
Main Methods:
- Utilized experimentally obtained PPI pairs as a training dataset.
- Extracted eight top-ranking keywords from Gene Ontology (GO) molecular function annotations.
- Developed two knowledge rules based on these keywords and protein co-localization to filter false positive PPIs.
- Defined a 'strength' metric (signal-to-noise ratio) to evaluate the effectiveness of the applied knowledge rules.
Main Results:
- The selected GO keywords demonstrated high sensitivity (64.21% in yeast, 80.83% in worm) for identifying true PPIs.
- The knowledge rules successfully reduced false positive PPIs, with the 'strength' of improvement ranging from two to ten-fold compared to random removal.
- Applied rules significantly improved the precision of predicted PPI datasets across different computational methods and organisms.
Conclusions:
- Gene Ontology annotations and deduced knowledge rules can effectively reduce false positives in computationally predicted PPI datasets.
- This approach enhances the true positive fraction and robustness of predicted PPIs, leading to better agreement with experimental findings.
- The developed method offers a valuable strategy for improving the reliability of large-scale PPI data for biological research.
Related Concept Videos
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.

