Related Experiment Video
Updated: Aug 6, 2026

Characterization of a Pathogenic Escherichia coli Strain Derived from Oreochromis spp. Farms Using Whole-Genome Sequencing
Published on: December 23, 2022
SEVA: structural and evolutionary feature integration for predicting virulence factors and antibiotic resistance
Kaiqi Li1, Xin Peng2, Xiuwei Qian1
1Department of Computer Science, City University of Hong Kong, Hong Kong SAR, 999077, China.
Background:
Infectious diseases continue to pose unprecedented challenges to public health and the global economy. Virulence factors (VFs) enable pathogens to adhere, reproduce, and cause damage to host cells, while antibiotic resistance genes (ARGs) enable pathogens to withstand treatments that would otherwise be effective. The concurrent identification of VFs and ARGs is crucial for efficient pathogen surveillance. However, existing tools for predicting VFs or ARGs typically suffer from high false negative rates and limitations in identifying only high-identity genes against known reference VF or ARG databases.
Results:
To address these challenges, we developed SEVA, an advanced model that integrates protein language models (pLMs) with structural and evolutionary protein features to predict VFs and ARGs from genome sequencing data. Integrating multiple homologous sequences can identify latent virulence or drug resistance caused by site mutations, reducing false negative rates. Meanwhile, the protein structure remains conserved despite the low sequence identity in some functional domains of VFs or ARGs. The aggregate of protein structure information further improves the identification abilities of VF and ARG. In addition, pLMs enable the model to capture high-dimensional feature representations more effectively. SEVA rigorously collected three datasets with over 20,000 genes and five reference databases. It outperforms state-of-the-art methods, including Diamond, VRprofile, FoldSeek, PreVFs-RG, PLM-ARG, ARG-BERT, and HyperVR, achieving an accuracy of 97.13% and confirming the efficacy of its key components, such as refined feature selection and multiple sequence alignment subsampling.
Conclusion:
SEVA takes protein sequences as input and derives evolutionary, structural, and statistical representations for prediction, making our model a reliable tool for VF and ARG prediction. This capability is particularly valuable in epidemic prevention and control, where accurate identification of VFs and ARGs is crucial. By providing concurrent and reliable predictions of VFs and ARGs, SEVA enhances our ability to respond to microbial threats effectively. This finding supports robust efforts to mitigate the spread of infectious diseases and safeguard public health, addressing a critical gap in contemporary epidemic response strategies. The SEVA model and data are available at https://github.com/kaiqili2/SEVA. Video Abstract.
Related Concept Videos
Regulation of Bacterial Virulence
Determinants of Bacterial Pathogenicity and Virulence
Modern Molecular Taxonomy
Rapid Identification of Pathogens
Mechanism of Antibiotic Resistance in MRSA
Development of Antibiotic Resistance
