Related Experiment Video
Updated: Oct 15, 2025

06:50
Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
2.1K
Protein-Protein Interaction Sites Prediction Based on an Under-Sampling Strategy and Random Forest Algorithm
IEEE/ACM Transactions on Computational Biology and Bioinformatics
|October 27, 2021
Summary
This study introduces a new computational method, NearMiss-based under-sampling for unbalancing datasets and Random Forest classification (NM-RF), to improve protein-protein interaction site prediction. The NM-RF approach effectively addresses class imbalance, enhancing prediction accuracy for interface residues.
Area of Science:
- Computational Biology
- Bioinformatics
- Structural Biology
Background:
- Protein-protein interactions (PPIs) are crucial for cellular functions.
- Predicting PPI sites computationally offers an alternative to costly experimental methods.
- Class imbalance between interface and non-interface residues hinders prediction accuracy.
Purpose of the Study:
- To develop a novel computational strategy for predicting protein-protein interaction sites.
- To address the challenge of class imbalance in protein sequence datasets.
- To enhance the accuracy of differentiating interface from non-interface residues.
Main Methods:
- Utilized Position-Specific Scoring Matrix (PSSM)-derived features, hydropathy index (HI), and relative solvent accessibility (RSA) for residue representation.
- Implemented NearMiss-based under-sampling to balance the dataset by reducing non-interface residues.
- Employed the Random Forest (RF) algorithm for binary classification of residues.
Main Results:
- The proposed NearMiss-based under-sampling for unbalancing datasets and Random Forest classification (NM-RF) model achieved high prediction accuracy.
- Achieved 87.6% accuracy on the Dtestset72 dataset.
- Achieved 84.3% accuracy on the PDBtestset164 dataset.
Conclusions:
- The NM-RF method effectively overcomes class imbalance issues in predicting protein interaction sites.
- The developed method demonstrates significant improvement in differentiating interface and non-interface residues.
- NM-RF offers a promising computational tool for advancing PPI site prediction research.
Related Concept Videos
Protein-protein Interfaces
14.0K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.0K
Conserved Binding Sites
4.7K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.7K
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Ligand Binding Sites
14.2K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
14.2K
Protein-Protein Interfaces
4.0K
4.0K
Conservation of Protein Domains Over Different Proteins
13.2K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
13.2K

