Related Experiment Video
Updated: Nov 29, 2025

06:50
Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
2.3K
Protein-protein interaction site prediction using random forest proximity distance.
Zhijun Qiu1,2, Qingjie Liu1
1College of Food and Bioengineering, Henan University of Science and Technology, Luoyang, P. R. China.
Journal of Bioinformatics and Computational Biology
|November 20, 2020
Summary
This study introduces a proximity distance (PD) method using random forests to enhance protein-protein interaction site (PPIS) prediction. Optimized PD metrics significantly improved prediction accuracy, demonstrating PD
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning in Biology
Background:
- Accurate prediction of protein-protein interaction sites (PPIS) is crucial for understanding cellular mechanisms.
- Existing distance metrics for PPIS prediction have limitations in classification accuracy.
- Random forest proximity distance (PD) offers a novel approach for PPIS prediction.
Purpose of the Study:
- To develop and evaluate a front-end method using random forest proximity distance (PD) for improved protein-protein interaction site (PPIS) prediction.
- To optimize the PD metric through an iterative method to enhance classification performance.
- To assess the reliability of PD in indicating prediction accuracy.
Main Methods:
- A front-end method utilizing random forest proximity distance (PD) was employed for PPIS prediction.
- Numerical analysis and statistical inference were used to compare PD with Mahalanobis and Cosine distances.
- An iterative method was designed to optimize PD by adjusting the random forest model's training set size, yielding 75PD and 50PD metrics.
- Performance was evaluated on two independent test sets using Matthews correlation coefficient and F1 score.
Main Results:
- Proximity distance (PD) demonstrated superior performance compared to Mahalanobis and Cosine distances in PPIS prediction.
- The iterative method yielded optimized PD metrics (75PD and 50PD) that significantly improved Matthews correlation coefficient and F1 scores on independent test sets.
- A statistically significant improvement in prediction accuracy was observed with the optimized PD metrics.
- Results indicated a positive correlation between the proximity of test data to training data and prediction accuracy.
Conclusions:
- The iterative method effectively optimizes proximity distance definitions for enhanced PPIS prediction.
- Proximity distance (PD) serves as a reliable indicator of prediction result reliability.
- The developed PD-based method offers a promising advancement in predicting protein-protein interaction sites.
Related Concept Videos
Protein-protein Interfaces
14.2K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.2K
Protein-Protein Interfaces
4.2K
4.2K
Protein Networks
4.4K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.4K
Protein Networks
2.6K
2.6K
Conserved Binding Sites
4.9K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.9K
Ligand Binding Sites
14.5K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
14.5K

