Related Experiment Video
Updated: Jul 21, 2025

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
CSI: Contrastive data Stratification for Interaction prediction and its application to compound-protein interaction
Apurva Kalia1, Dilip Krishnan2, Soha Hassoun1,3
1Department of Computer Science, Tufts University, Medford, MA 02155, United States.
We introduce Contrastive Stratification for Interaction Prediction (CSI), a novel method for partitioning data to improve interaction prediction. CSI enhances deep learning models by creating multi-views for contrastive learning, significantly boosting prediction accuracy in areas like drug discovery.
Area of Science:
- Computational Biology
- Machine Learning
- Bioinformatics
Background:
- Accurate prediction of interactions between biological entities (e.g., compound-protein) is crucial for drug discovery and synthetic biology.
- Current deep learning models often struggle to fully leverage the relational information inherent in interaction datasets.
- Exploiting multi-view representations of interacting objects can enhance model performance through contrastive learning.
Purpose of the Study:
- To develop a novel method, Contrastive Stratification for Interaction Prediction (CSI), for partitioning interaction datasets.
- To improve the learning of object representations by utilizing congruent and non-congruent data views via contrastive learning.
- To apply CSI to the compound-protein interaction prediction problem to accelerate drug discovery and related applications.
Main Methods:
- CSI stratifies (partitions) datasets by assigning a key and multiple views to each data point.
- Data partitions under a specific key form congruent views, enabling contrastive multiview coding.
- The method learns embeddings that maximize mutual information across these congruent views.
Main Results:
- CSI significantly improved average precision in compound-protein interaction prediction, with gains ranging from 13.7% to 39% when using compounds/sequences as keys.
- Further gains of 16.9% to 63% were observed when using reaction features as keys in enzymatic datasets.
- These results demonstrate the effectiveness of data stratification and contrastive learning for interaction prediction.
Conclusions:
- CSI offers a powerful approach to enhance interaction prediction by effectively leveraging multi-view data representations.
- The method shows substantial improvements over baseline models lacking data stratification and contrastive learning.
- CSI has the potential to expedite drug discovery, metabolic engineering, and synthetic biology applications.
More Related Videos
Related Concept Videos
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein-Protein Interfaces
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...

