Related Experiment Video
Updated: May 11, 2026

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
Published on: January 26, 2024
Semi-supervised prediction of SH2-peptide interactions from imbalanced high-throughput data
Kousik Kundu1, Fabrizio Costa, Michael Huber
1Bioinformatics Group, Department of Computer Science, University of Freiburg, Freiburg, Germany.
We developed a machine learning method to predict interactions between human Src homology 2 (SH2) domains and phosphotyrosine peptides, improving accuracy over existing methods. Our approach enhances understanding of cellular processes by predicting SH2-peptide binding partners.
Area of Science:
- Biochemistry
- Computational Biology
- Bioinformatics
Background:
- Src homology 2 (SH2) domains are crucial peptide-recognition modules binding phosphotyrosine peptides.
- Understanding SH2-domain binding partners is key to deciphering cellular processes.
- Existing in-silico prediction methods for SH2-peptide interactions have limitations in coverage, modeling, and computational complexity.
Purpose of the Study:
- To develop a novel, effective machine learning approach for predicting human SH2-peptide interactions.
- To improve upon the performance of existing state-of-the-art prediction methods.
- To provide a valuable tool for the scientific community to predict SH2-peptide binding.
Main Methods:
- Utilized comprehensive data from micro-array and peptide-array experiments for 51 human SH2 domains.
- Employed a semi-supervised machine learning setting to address data imbalance and high signal-to-noise ratio.
- Incorporated information from non-interacting peptides (negative examples) and considered high-order correlations using regularization techniques.
Main Results:
- Achieved a high predictive performance with 0.83 AUC ROC and 0.93 AUC PR, surpassing the PSSM-based SMALI approach (0.71 AUC ROC, 0.87 AUC PR).
- Demonstrated the benefit of including negative examples and high-order correlations in modeling.
- Successfully tackled data imbalance using a semi-supervised strategy.
Conclusions:
- The developed machine learning approach offers a significant improvement for predicting SH2-peptide interactions.
- The study provides valuable insights into SH2-domain binding specificities and their biological relevance through genome-wide predictions.
- Models and genome-wide predictions are made publicly available to facilitate further research.
More Related Videos
Related Concept Videos
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Protein-protein Interfaces

