Related Experiment Video
Updated: Jun 16, 2026

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
Published on: January 26, 2024
Genome-wide sequence-based prediction of peripheral proteins using a novel semi-supervised learning technique
Nitin Bhardwaj1, Mark Gerstein, Hui Lu
1Bioinformatics Program, Department of Bioengineering, University of Illinois at Chicago, Chicago, IL 60607, USA. nitin.bhardwaj@yale.edu
This study introduces a novel positive-unlabeled (PU) learning approach for identifying membrane-binding protein domains. The developed protocol accurately predicts these crucial biological components, achieving up to 95% accuracy.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning
Background:
- Traditional supervised learning requires labeled data from all classes, which is often infeasible for biological datasets.
- Peripheral domains, crucial for cell signaling and trafficking, present a challenge due to the lack of comprehensive negative binding data.
- Positive-unlabeled (PU) learning offers a solution by utilizing a set of known positive examples and a larger set of unlabeled examples.
Purpose of the Study:
- To apply PU learning for the first time to predict protein function, specifically identifying membrane-binding peripheral domains.
- To develop and validate a computational protocol for predicting domain-membrane interactions using limited labeled data.
Main Methods:
- Implemented an iterative PU learning algorithm to identify reliable negative examples from the unlabeled dataset.
- Constructed a classifier using the positive set and the identified reliable negative set.
- Utilized a dataset comprising 232 positive cases and approximately 3750 unlabeled cases for protocol development and validation.
Main Results:
- The developed PU learning protocol achieved a prediction accuracy of up to 95% in holdout evaluations.
- Independent implementations confirmed the robustness and high accuracy of the prediction protocol.
- The method effectively addresses the challenge of predicting membrane-binding properties when only positive examples are readily available.
Conclusions:
- The study demonstrates the efficacy of PU learning for predicting membrane-binding properties of protein domains.
- The developed protocol is valuable for biological research, particularly in cases where data is skewed towards one class.
- This approach enhances the ability to identify and study functionally important protein domains involved in cellular processes.
More Related Videos
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Evolutionary Relationships through Genome Comparisons
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

