Related Experiment Video
Updated: Oct 9, 2026

Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
Targeted genome recovery of under-sequenced microbes from the Sequence Read Archive STAT
Harold P Hodgins1,2, Briallen Lobb1,2, Mackenzie Peck1
1Department of Biology, University of Waterloo, 200 University Avenue West, Waterloo, ON, N2L 3G1, Canada.
Abstract:
Most microbial species are represented by a single genome in public databases. The lack of genomes for these 'singleton' organisms limits our understanding of their pan-genome diversity and evolution. Although the Sequence Read Archive (SRA) contains millions of sequencing datasets that could be used to expand our understanding of many species, the extent to which under-represented microbial species are present at levels sufficient for genome recovery is unclear. Here, we show that the pre-computed taxonomic profiles generated by the National Center for Biotechnology Information SRA Taxonomy Analysis Tool (STAT) can be used to identify SRA datasets containing recoverable genomes for under-represented microbes. Across >28 million SRA datasets, tens of thousands of singleton archaeal and bacterial species were detected, often at abundances consistent with successful genome recovery. Applying targeted genome recovery to 804 singleton species, we recovered genomes representing new strains for 472 species. The success rate of genome recovery correlated with STAT-derived estimates of genome coverage, demonstrating that genome recovery from the SRA is both predictable and scalable. Using the single available genome of Clostridium tarantellae as a case study, SRA data mining recovered seven additional C. tarantellae genomes, correcting assembly gaps in the reference genome, expanding its pan-genome and increasing its known host range by seven additional fish species. These findings reveal that many microbial species currently represented by a single genome are in fact widely distributed across existing sequencing data and highlight a major opportunity to systematically expand strain-level genomic diversity and pan-genomic representation for under-sampled microbial species without additional sequencing.

