MultiStageSearch: An Iterative Workflow for Unbiased Taxonomic Analysis of Pathogens Using Proteogenomics
Julian Pipart1, Tanja Holstein1,2,3,4,5, Lennart Martens2,3,4,5
1Data Competence Center MF 2, Robert Koch Institute, Berlin 13353, Germany.
None:
The global SARS-CoV-2 pandemic emphasized the need for accurate pathogen diagnostics. While genomics is the gold standard, integrating mass spectrometry-based proteomics offers additional benefits. However, current proteomic and genomic reference databases are often biased toward specific taxa, such as pathogenic strains or model organisms, and proteomic databases are less comprehensive. These biases and gaps can lead to inaccurate identifications. To address these issues, we introduce MultiStageSearch, a multistep database search method that combines proteome and genome databases for taxonomic analysis. Initially, a generalist proteome database is used to infer potential species. Then, MultiStageSearch generates a specialized proteogenomic database for precise identification. This database is preprocessed to filter duplicates and cluster identical open reading frames to reduce genomic database biases. The workflow operates independently of strain-level NCBI taxonomy, enabling the identification of strains not represented in existing taxonomies. We benchmarked the workflow on viral and bacterial samples, demonstrating its superior performance in strain-level taxonomic inference compared to existing methods. MultiStageSearch offers a flexible and accurate approach for pathogen research and diagnostics, overcoming incomplete search spaces and biases inherent in reference databases.
More Related Videos
05:37Label-Free Quantitative Proteomics Workflow for Discovery-Driven Host-Pathogen Interactions
Published on: October 20, 2020
08:09An Aquatic Microbial Metaproteomics Workflow: From Cells to Tryptic Peptides Suitable for Tandem Mass Spectrometry-based Analysis
Published on: September 15, 2015
