Related Experiment Video
Updated: Jul 1, 2025

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Cracking the black box of deep sequence-based protein-protein interaction prediction
Judith Bernett1, David B Blumenthal2, Markus List1
1Data Science in Systems Biology, TUM School of Life Sciences, Technical University of Munich, Maximus-von-Imhof Forum 3, 85354, Freising, Germany.
Deep learning models for protein-protein interaction (PPI) prediction often overestimate performance due to data leakage. When leakage is avoided, performance drops, highlighting the need for better computational methods and experimental research.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning
Background:
- Protein-protein interactions (PPIs) are fundamental to biological processes.
- Computational methods are developed to predict PPIs as cost-effective alternatives to experiments.
- Existing prediction methods report high accuracy, but reproducibility is often questionable.
Purpose of the Study:
- To systematically evaluate the reproducibility of deep learning models for PPI prediction.
- To investigate the impact of data leakage, sequence similarity, and node degree on model performance.
- To compare deep learning models against baseline machine learning approaches.
Main Methods:
- Deep learning models and basic machine learning models were employed for PPI prediction.
- The study involved systematic investigation of data leakage through set overlaps.
- Sequence similarity and node degree information were analyzed as key features.
- Performance was evaluated under conditions with and without data leakage.
Main Results:
- Random splitting of training and test sets leads to significant overestimation of model performance due to data leakage.
- When data leakage is prevented by minimizing sequence similarity, deep learning model performance becomes comparable to random chance.
- Baseline models utilizing sequence similarity and network topology achieve good performance with lower computational cost.
- Deep learning models primarily learn from sequence similarities and node degrees when data leakage is present.
Conclusions:
- Overestimated performance in PPI prediction is a common issue, often stemming from data leakage.
- Current deep learning approaches may not offer substantial improvements over simpler baseline methods when data leakage is controlled.
- Predicting PPIs for proteins with low sequence similarity to known proteins remains a significant challenge.
- Further experimental investigation of the 'dark' interactome and development of robust computational methods are essential.
Related Concept Videos
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein-Protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...

