A hybrid deep learning framework for WT or mutant peptide prediction using p53 mutation data
Manisha R Patil1, Anand Bihari2
1School of Computer Science Engineering and Information System, Vellore Institute of Technology, Vellore, Tamil Nadu, India.
Abstract:
p53 is a tumor suppressor protein that maintains genome integrity. Single amino acid variations (SAVs), especially hotspot mutations (such as R175 and R248), are closely associated with oncogenic transformation and weaken DNA-binding and transcriptional regulatory capabilities. To identify deleterious SAVs for understanding disease mechanisms and key cellular processes, including apoptosis, DNA repair, and cell cycle regulation. This study used quantitative biochemical and microbial descriptors of p53 peptide sequences, along with a CNN and a 2-layer bidirectional long short-term memory (Bi-LSTM) network architecture with attention mechanisms and ESM-2-based embeddings, to classify sequences as wild-type or mutant. A systematic analysis of the p53 mutation dataset was performed using molecular weight, instability index, hydrophobicity, motif enrichment, and amino acid substitution patterns to identify mutation-centered peptides. Using stratified 5-fold cross-validation, the proposed model was tested, achieving an accuracy of 0.98, an area under the ROC curve (AUROC) of 0.98, and a precision-recall AUC of 0.99. Furthermore, SHAP-based interpretation identified the key amino acid residues and biochemical factors that contribute to the model's predictive performance. The proposed approach provides insights into the structural, sequential, and biochemical effects of variants, an interpretable, robust framework for evaluating the functional consequences of p53 hotspot and other variants, and a computationally efficient tool for prioritizing high-risk p53 variants.
Insights
This study developed a deep learning model to identify harmful p53 protein mutations. The model accurately predicts the impact of single amino acid variations (SAVs) on tumor suppression, aiding cancer research.
Area of Science:
- Molecular Biology
- Bioinformatics
- Computational Biology
Background:
- The p53 protein is crucial for maintaining genome integrity and acts as a tumor suppressor.
- Single amino acid variations (SAVs), particularly hotspot mutations, can disrupt p53's DNA-binding and transcriptional functions, promoting oncogenesis.
- Understanding the functional impact of p53 SAVs is vital for elucidating disease mechanisms and cellular processes like apoptosis and DNA repair.
Purpose of the Study:
- To develop a computational framework for identifying deleterious p53 SAVs.
- To analyze the structural, sequential, and biochemical effects of p53 variants.
- To provide an interpretable and efficient tool for prioritizing high-risk p53 mutations.
Main Methods:
- Utilized quantitative biochemical descriptors and sequence data of p53 peptides.
- Employed a deep learning architecture combining Convolutional Neural Networks (CNN) and bidirectional Long Short-Term Memory (Bi-LSTM) with attention mechanisms and ESM-2 embeddings.
- Performed systematic analysis using molecular weight, instability index, hydrophobicity, motif enrichment, and amino acid substitution patterns.
Main Results:
- Achieved high predictive performance with 0.98 accuracy, 0.98 AUROC, and 0.99 precision-recall AUC using stratified 5-fold cross-validation.
- Identified key amino acid residues and biochemical factors influencing p53 variant pathogenicity via SHAP-based interpretation.
- Demonstrated the model's robustness in classifying wild-type versus mutant p53 sequences.
Conclusions:
- The developed model offers a robust and interpretable framework for assessing the functional consequences of p53 variants.
- Provides insights into the biochemical and structural impacts of SAVs on p53.
- Presents a computationally efficient tool for prioritizing p53 variants with high oncogenic risk.
More Related Videos
04:56Detection of Aggregation-Prone Behavior in Mutant P53 V157F Breast Cancer Cells Using Multipoint Thioflavin T Fluorescence
Published on: December 30, 2025
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
