Related Experiment Video
Updated: Feb 9, 2026

Capsular Serotyping of Streptococcus pneumoniae by Latex Agglutination
Published on: September 25, 2014
SeroBA: rapid high-throughput serotyping of Streptococcus pneumoniae from whole genome sequence data
Lennard Epping1,2, Andries J van Tonder3, Rebecca A Gladstone3
12Microbial Genomics, Robert Koch Institute, Berlin, Germany.
Insights
A new k-mer based method, SeroBA, accurately identifies Streptococcus pneumoniae serotypes directly from whole genome sequencing data. This scalable approach aids in tracking serotype evolution post-vaccine introduction.
Area of Science:
- Microbiology
- Genomics
- Bioinformatics
Background:
- Streptococcus pneumoniae causes significant child mortality globally.
- Accurate serotyping is crucial for monitoring vaccine impact and pathogen evolution.
- Existing genomic serotyping tools have limitations in scalability and performance.
Purpose of the Study:
- To introduce SeroBA, a novel k-mer based method for rapid and accurate pneumococcal serotype identification.
- To evaluate SeroBA's performance, scalability, and robustness using real and simulated genomic data.
- To enable efficient serotype prediction directly from raw whole genome sequencing reads.
Main Methods:
- Development of SeroBA, a k-mer based bioinformatics tool.
- Comparative analysis of SeroBA against existing methods using validation and large datasets.
- Assessment of SeroBA's performance across varying whole genome sequencing coverage depths.
Main Results:
- SeroBA achieves 98% concordance in serotype prediction by identifying the cps locus.
- The method demonstrates high scalability, processing 10,000 samples in approximately one day.
- Accurate serotyping is achievable with low sequence coverage (15-21×).
Conclusions:
- SeroBA offers a robust, scalable, and accurate solution for pneumococcal serotype identification from WGS data.
- The tool facilitates efficient genomic epidemiology and surveillance of Streptococcus pneumoniae.
- SeroBA is freely available as an open-source Python3 package.
Abstract:
Streptococcus pneumoniae is responsible for 240 000-460 000 deaths in children under 5 years of age each year. Accurate identification of pneumococcal serotypes is important for tracking the distribution and evolution of serotypes following the introduction of effective vaccines. Recent efforts have been made to infer serotypes directly from genomic data but current software approaches are limited and do not scale well. Here, we introduce a novel method, SeroBA, which uses a k-mer approach. We compare SeroBA against real and simulated data and present results on the concordance and computational performance against a validation dataset, the robustness and scalability when analysing a large dataset, and the impact of varying the depth of coverage on sequence-based serotyping. SeroBA can predict serotypes, by identifying the cps locus, directly from raw whole genome sequencing read data with 98 % concordance using a k-mer-based method, can process 10 000 samples in just over 1 day using a standard server and can call serotypes at a coverage as low as 15-21×. SeroBA is implemented in Python3 and is freely available under an open source GPLv3 licence from: https://github.com/sanger-pathogens/seroba.
Related Concept Videos
Genomics
Genomic Imprinting and Inheritance
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
Pneumonia I: Introduction
Risk Factors
Various factors influence the likelihood of developing pneumonia. Age plays a crucial role, with infants, children under two, and individuals over 65 at increased risk due to their...
Pneumonia II: Pathophysiology
Genome Size and the Evolution of New Genes
Cis-regulatory Sequences

