Advancing the Accuracy of Anti-MRSA Peptide Prediction Through Integrating Multi-Source Protein Language Models

Watshara Shoombuatong1, Pakpoom Mookdarsanit2, Lawankorn Mookdarsanit3

  • 1Center for Research Innovation and Biomedical Informatics, Faculty of Medical Technology, Mahidol University, Bangkok, 10700, Thailand. watshara.sho@mahidol.ac.th.

Insights

A new computational framework, pLM4MRSA, accurately predicts anti-methicillin-resistant Staphylococcus aureus (MRSA) peptides. This method uses pre-trained protein language models (pLMs) and principal component analysis (PCA) to improve drug discovery efficiency.

Area of Science:

  • Computational biology
  • Drug discovery
  • Bioinformatics

Background:

  • Methicillin-resistant Staphylococcus aureus (MRSA) infections necessitate novel therapeutic strategies.
  • Current experimental methods for identifying anti-MRSA peptides are time-consuming and resource-intensive.
  • Computational approaches are needed for efficient, sequence-based prediction of anti-MRSA peptides.

Purpose of the Study:

  • To develop a highly accurate computational framework, pLM4MRSA, for predicting anti-MRSA peptides.
  • To leverage pre-trained protein language models (pLMs) for enhanced feature extraction from peptide sequences.
  • To improve the efficiency and accuracy of identifying potential anti-MRSA drug candidates.

Main Methods:

  • Utilized pre-trained protein language models (pLMs) like ProtTrans and ESM-2 to generate hybrid feature embeddings.
  • Applied principal component analysis (PCA) to reduce dimensionality and process the hybrid embeddings.
  • Constructed a predictive model using PCA-transformed feature vectors for anti-MRSA peptide identification.

Main Results:

  • The pLM4MRSA framework achieved a balanced accuracy of 0.983 and a Matthew correlation coefficient of 0.980 on an independent test dataset.
  • Demonstrated significant improvements over existing state-of-the-art methods, with accuracy gains of 2.53%-4.83% and MCC gains of 7.73%-13.23%.
  • Hybrid feature embeddings effectively captured discriminative patterns, outperforming traditional hand-crafted features.

Conclusions:

  • pLM4MRSA offers a high-performance, accurate, and efficient computational solution for predicting anti-MRSA peptides.
  • The framework shows excellent applicability in drug discovery pipelines for identifying novel anti-MRSA agents.
  • The integration of diverse pLM embeddings provides a powerful approach for in silico peptide property prediction.