Related Experiment Video
Updated: Sep 15, 2025

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Deep learning models for unbiased sequence-based PPI prediction plateau at an accuracy of 0.65.
Timo Reim1,2, Anne Hartebrodt2, David B Blumenthal2
1Data Science in Systems Biology, TUM School of Life Sciences, Technical University of Munich, Freising, 85354, Germany.
Protein-protein interaction prediction using sequence-based methods is challenging due to data leakage. While ESM-2 embeddings improve performance, reliable prediction may require structural data.
Area of Science:
- Computational biology
- Bioinformatics
- Protein interaction analysis
Background:
- Protein-protein interactions (PPIs) are crucial for cellular functions.
- Accurate computational prediction of PPIs remains a significant challenge.
- Previous evaluation schemes and data leakage issues have obscured progress in sequence-based PPI prediction.
Purpose of the Study:
- To investigate the impact of protein embeddings on sequence-based PPI prediction.
- To evaluate the performance of different model architectures and embedding strategies.
- To determine if sequence-based models can implicitly learn contact maps for PPI prediction.
Main Methods:
- Utilized ESM-2 protein embeddings for sequence-based PPI prediction.
- Compared models with varying complexity, per-protein, and per-token embeddings.
- Assessed the influence of self-attention and cross-attention mechanisms.
- Analyzed the ability of models to learn contact maps as intermediate representations.
Main Results:
- ESM-2 embeddings significantly explain performance gains in PPI prediction, irrespective of model architecture.
- All tested sequence-based models plateaued at an accuracy of 0.65.
- Sequence-based models cannot implicitly learn a contact map.
- Performance gains are attributed to embeddings rather than inherent sequence-based learning capabilities.
Conclusions:
- Protein embeddings like ESM-2 are key drivers of recent performance improvements in sequence-based PPI prediction.
- Current sequence-based models have limitations and cannot implicitly learn contact maps.
- Structural information may be essential for achieving reliable and accurate PPI predictions.
More Related Videos
Related Concept Videos
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein-Protein Interfaces
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Folding Quality Check in the RER

