Related Experiment Video
Updated: Aug 13, 2026

08:38
Genome-wide Protein-protein Interaction Screening by Protein-fragment Complementation Assay (PCA) in Living Cells
Published on: March 3, 2015
Evaluation of different biological data and computational classification methods for use in protein interaction
Yanjun Qi1, Ziv Bar-Joseph, Judith Klein-Seetharaman
1School of Computer Science, Carnegie Mellon University, Pittsburgh, Pennsylvania 15213, USA.
Proteins
|February 2, 2006
Summary
Predicting protein interactions is crucial for understanding biological systems. Gene expression data proved most valuable across all prediction tasks, outperforming yeast-2-hybrid methods for supervised learning models.
Area of Science:
- Bioinformatics
- Computational Biology
- Systems Biology
Background:
- Protein-protein interactions (PPIs) are fundamental to biological processes.
- High-throughput methods for PPI detection in yeast are often incomplete and error-prone.
- Supervised learning offers a promising approach to integrate diverse biological data for PPI prediction.
Purpose of the Study:
- To systematically evaluate the utility of various biological data sources and feature encodings for PPI prediction.
- To compare the performance of different classifiers across three distinct PPI tasks: physical interaction, co-complex, and pathway co-membership.
- To identify the most influential features for accurate PPI prediction.
Main Methods:
- Assembled a comprehensive set of biological features and varied their encoding.
- Employed six classifiers: Random Forest (RF), RF k-NN, Naïve Bayes, Decision Tree, Logistic Regression, and Support Vector Machine.
- Utilized RF's Gini index for feature importance assessment and evaluated accuracy with top-ranked features.
Main Results:
- Co-complex relationship prediction was generally easier than physical interaction or pathway co-membership.
- The Random Forest classifier consistently performed as a top-tier model across all feature sets and tasks.
- Gene expression emerged as the most critical feature for all prediction tasks, irrespective of encoding.
- Yeast-2-hybrid data was not a top-ranking feature under any tested condition.
Conclusions:
- Supervised learning, particularly with Random Forest, is effective for PPI prediction.
- Feature importance is task-dependent and influenced by data encoding.
- Gene expression data holds significant predictive power for various types of protein interactions.
More Related Videos
Related Concept Videos
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein-Protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Ligand Binding Sites
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...

