Related Experiment Video
Updated: Jun 28, 2025

06:50
Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
1.8K
Protein feature engineering framework for AMPylation site prediction.
Hardik Prabhu1,2, Hrushikesh Bhosale1, Aamod Sane1
1Computing and Data Sciences, FLAME University, Pune, 412115, India.
Scientific Reports
|April 15, 2024
Summary
This study introduces a new computational method for identifying AMPylation sites, a crucial protein modification. Our novel feature extraction pipeline improves machine learning model accuracy for predicting these sites.
Area of Science:
- Biochemistry
- Bioinformatics
- Computational Biology
Background:
- AMPylation is a significant post-translational modification involving the addition of adenosine monophosphate (AMP) to tyrosine and threonine residues.
- While its prevalence and impact are increasingly understood, experimental identification of AMPylation sites is difficult.
- Computational prediction offers a faster alternative, but model performance relies heavily on feature representation.
Purpose of the Study:
- To develop a novel feature extraction pipeline for improved computational prediction of AMPylation sites.
- To evaluate the effectiveness of these novel features in machine learning models for AMPylation site identification.
Main Methods:
- A new feature extraction pipeline was designed to encode key properties relevant to AMPylation.
- Machine learning classifiers were trained using numerical representations derived from the extracted features.
- Tenfold cross-validation was employed to assess model performance in distinguishing AMPylated from non-AMPylated sites.
- SHapley Additive exPlanations (SHAP) were used to analyze feature importance.
Main Results:
- The developed feature extraction framework demonstrated utility in machine learning models.
- The top-performing features achieved a Matthews Correlation Coefficient (MCC) score of 0.58, Accuracy of 0.8, AUC-ROC of 0.85, and F1 score of 0.73.
- Analysis elucidated model behavior based on monogram and bigram counts using SHAP.
Conclusions:
- The novel feature extraction pipeline significantly enhances the predictive performance of machine learning models for AMPylation sites.
- This approach offers a more efficient and accurate method for identifying AMPylation sites compared to experimental methods.
- The findings contribute to a better understanding of AMPylation and facilitate further research into its biological roles.
Related Concept Videos
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Allosteric Proteins-ATCase
5.7K
Binding sites linkages can regulate a protein's function. For example, enzyme activity is often regulated through a feedback mechanism where the end product of the biochemical process serves as an inhibitor.
Aspartate transcarbamoylase (ATCase) is a cytosolic enzyme that catalyzes the condensation of L-aspartate and carbamoyl phosphate to N-carbamoyl-L-aspartate. This reaction is the first step in pyrimidine biosynthesis. UTP and CTP, the end products of the pyrimidine synthesis...
Aspartate transcarbamoylase (ATCase) is a cytosolic enzyme that catalyzes the condensation of L-aspartate and carbamoyl phosphate to N-carbamoyl-L-aspartate. This reaction is the first step in pyrimidine biosynthesis. UTP and CTP, the end products of the pyrimidine synthesis...
5.7K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Ligand Binding and Linkage
4.8K
Allosteric proteins have more than one ligand binding site; the binding of a ligand to any of these sites influences the binding of ligands to the other sites. When a protein is allosteric, its binding sites are called coupled or linked. In the case of enzymes, the site that binds to the substrate is known as the active site and the other site is known as the regulatory site. When a ligand binds to the regulatory site, this leads to conformational changes in the protein that can influence...
4.8K

