Related Experiment Video
Updated: Sep 2, 2025

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Prediction of protein mononucleotide binding sites using AlphaFold2 and machine learning
Shohei Yamaguchi1, Haruka Nakashima1, Yoshitaka Moriwaki2
1Department of Biotechnology, The University of Tokyo, Japan.
This study introduces a novel system for predicting protein binding sites for five essential mononucleotides. The ensemble predictor achieved the best performance, demonstrating high accuracy in identifying these crucial molecular interactions.
Area of Science:
- Biochemistry
- Computational Biology
- Structural Biology
Background:
- Protein-ligand interactions are fundamental to cellular processes.
- Accurate prediction of nucleotide-binding sites is crucial for understanding protein function and drug discovery.
- Existing methods often lack comprehensive accuracy for diverse nucleotide types.
Purpose of the Study:
- To develop and evaluate a robust system for predicting protein binding sites for five key mononucleotides: AMP, ADP, ATP, GDP, and GTP.
- To compare the performance of machine learning and template-based prediction methods.
- To create an ensemble model combining multiple prediction strategies for improved accuracy.
Main Methods:
- Developed two machine learning (ML) predictors (convolutional neural network, gradient boosting machine) utilizing protein structure features.
- Created two template-based predictors leveraging sequence and structure alignment.
- Employed data augmentation with similar ligand structures for enhanced ML model training.
- Utilized AlphaFold2 for generating accurate protein structure models.
- Constructed an ensemble predictor integrating the outputs of the four individual predictors.
Main Results:
- The template-based predictor using structure alignment demonstrated superior performance among individual methods.
- The ensemble predictor achieved the highest overall performance, yielding an Area Under the Curve (AUC) of 0.958 across all mononucleotides.
- Data augmentation and AlphaFold2-derived features improved ML predictor performance.
Conclusions:
- The developed ensemble system provides a highly accurate method for predicting mononucleotide-binding sites in proteins.
- Combining diverse prediction strategies (ML and template-based) is effective for enhancing prediction accuracy.
- This system has significant potential for applications in structural biology, drug discovery, and functional genomics.
More Related Videos
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-protein Interfaces
Ligand Binding and Linkage
Allosteric Proteins-ATCase
Aspartate transcarbamoylase (ATCase) is a cytosolic enzyme that catalyzes the condensation of L-aspartate and carbamoyl phosphate to N-carbamoyl-L-aspartate. This reaction is the first step in pyrimidine biosynthesis. UTP and CTP, the end products of the pyrimidine synthesis...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...

