Related Experiment Video
Updated: May 1, 2026

Biosensor for Detection of Antibiotic Resistant Staphylococcus Bacteria
Published on: May 8, 2013
A comparison of various feature extraction and machine learning methods for antimicrobial resistance prediction in
Deniz Ece Kaya1, Ege Ülgen1, Ayşe Sesin Kocagöz2
1Department of Biostatistics and Medical Informatics, School of Medicine, Acibadem Mehmet Ali Aydinlar University, Istanbul, Türkiye.
Abstract:
Streptococcus pneumoniae is one of the major concerns of clinicians and one of the global public health problems. This pathogen is associated with high morbidity and mortality rates and antimicrobial resistance (AMR). In the last few years, reduced genome sequencing costs have made it possible to explore more of the drug resistance of S. pneumoniae, and machine learning (ML) has become a popular tool for understanding, diagnosing, treating, and predicting these phenotypes. Nucleotide k-mers, amino acid k-mers, single nucleotide polymorphisms (SNPs), and combinations of these features have rich genetic information in whole-genome sequencing. This study compares different ML models for predicting AMR phenotype for S. pneumoniae. We compared nucleotide k-mers, amino acid k-mers, SNPs, and their combinations to predict AMR in S. pneumoniae for three antibiotics: Penicillin, Erythromycin, and Tetracycline. 980 pneumococcal strains were downloaded from the European Nucleotide Archive (ENA). Furthermore, we used and compared several machine learning methods to train the models, including random forests, support vector machines, stochastic gradient boosting, and extreme gradient boosting. In this study, we found that key features of the AMR prediction model setup and the choice of machine learning method affected the results. The approach can be applied here to further studies to improve AMR prediction accuracy and efficiency.
Insights
Machine learning models can predict antimicrobial resistance (AMR) in Streptococcus pneumoniae using genetic data. Different machine learning methods and feature types significantly impact prediction accuracy for antibiotics like penicillin.
Area of Science:
- Genomics and Bioinformatics
- Computational Biology
- Infectious Diseases
Background:
- Streptococcus pneumoniae poses a significant global health challenge due to high morbidity, mortality, and increasing antimicrobial resistance (AMR).
- Advances in whole-genome sequencing and machine learning (ML) offer new avenues for understanding and predicting AMR phenotypes in S. pneumoniae.
Purpose of the Study:
- To compare the effectiveness of different machine learning models and genetic features for predicting AMR in S. pneumoniae.
- To evaluate prediction accuracy for resistance to Penicillin, Erythromycin, and Tetracycline.
Main Methods:
- Utilized whole-genome sequencing data from 980 S. pneumoniae strains obtained from the European Nucleotide Archive (ENA).
- Extracted and compared genetic features including nucleotide k-mers, amino acid k-mers, and single nucleotide polymorphisms (SNPs).
- Trained and compared various machine learning models: random forests, support vector machines, stochastic gradient boosting, and extreme gradient boosting.
Main Results:
- The choice of machine learning method and the specific genetic features used significantly influenced the accuracy of AMR prediction.
- Different feature sets (k-mers, SNPs, combinations) yielded varying performance levels across the tested antibiotics.
Conclusions:
- Machine learning approaches are viable for predicting S. pneumoniae AMR phenotypes.
- Optimizing model setup and feature selection is crucial for enhancing prediction accuracy and efficiency in future AMR surveillance and clinical applications.
Related Concept Videos
Modern Molecular Taxonomy
Mechanism of Antibiotic Resistance in MRSA
Clinical Significance of Antibiotic Resistance

