Related Experiment Video
Updated: Jul 12, 2026

Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
PLM-VF: A novel predictor for virulent proteins in bacterial pathogens based on multi-scale CNN and protein language
Qingyang Guo1, Yusen Su1, Taigang Liu1
1College of Information Technology, Shanghai Ocean University, Shanghai 201306, China.
Abstract:
The rapid emergence and evolution of infectious diseases highlight the imperative need to detect virulence factors (VFs), which, in this study, specifically refer to virulent proteins in bacterial pathogens that can enhance pathogens' ability to infect and damage host cells. Traditional experimental approaches for the identification of VFs are labor-intensive and expensive, necessitating the development of more efficient computational techniques. In this study, we introduce a robust and interpretable model, named PLM-VF, which integrates embeddings from two pre-trained protein language models (PPLMs), i.e., ESM-1b and ProtT5, and utilizes Multi-Scale convolutional neural network (CNN) architectures to enhance the prediction accuracy (ACC) of VFs. In our evaluations, PLM-VF demonstrates strong performance on the independent dataset test, achieving the ACC of 86.98 %, the area under the receiver operating characteristic (ROC) curve (AUROC) of 0.9406, and the Matthews correlation coefficient (MCC) of 0.7442. These metrics underscore its superior ability in accurately identifying VFs based on their sequences, surpassing most of existing advanced methodologies. Furthermore, the proposed model has the capability of providing detailed insights into its decision-making process, which could support ongoing biological research and enhance potential clinical applications.
Insights
This study introduces PLM-VF, a new computational model for identifying bacterial virulent proteins (VFs). It accurately predicts VFs from protein sequences, aiding in infectious disease research and clinical applications.
Area of Science:
- Microbiology
- Bioinformatics
- Computational Biology
Background:
- Infectious diseases necessitate rapid detection of bacterial virulence factors (VFs).
- Traditional experimental VF identification is time-consuming and costly.
- Efficient computational methods are crucial for VF detection.
Purpose of the Study:
- To develop a robust and interpretable computational model for identifying bacterial virulent proteins (VFs).
- To improve the accuracy and efficiency of VF detection using protein sequences.
Main Methods:
- Integration of embeddings from two pre-trained protein language models (PPLMs): ESM-1b and ProtT5.
- Utilized Multi-Scale convolutional neural network (CNN) architectures for enhanced prediction.
- Model named PLM-VF developed for VF identification.
Main Results:
- PLM-VF achieved high performance on an independent test dataset.
- Accuracy (ACC) of 86.98%, AUROC of 0.9406, and MCC of 0.7442.
- Demonstrated superior VF identification capabilities based on protein sequences compared to existing methods.
Conclusions:
- PLM-VF offers a powerful and interpretable approach for identifying bacterial VFs.
- The model's performance surpasses current advanced methodologies.
- Potential applications in biological research and clinical settings for infectious disease management.

