Related Experiment Video
Updated: Jun 18, 2026

Efficient Nucleic Acid Extraction and 16S rRNA Gene Sequencing for Bacterial Community Characterization
Published on: April 14, 2016
Genestrip: exact and efficient read classification for selected groups of organisms
Daniel Pfeifer1, Markus Graf2, Clas Rurik3
1IT Faculty, Heilbronn University, Max-Planck-Str. 39, 74081, Heilbronn, Baden-Württemberg, Germany. daniel.pfeifer@hs-heilbronn.de.
Background:
The consumption of main memory resources is a significant burden in k-mer-based metagenomic analysis when creating related databases but also when performing (unique) k-mer-counting and read classification. Genestrip addresses this issue by focusing on small but freely configurable groups of organisms. Regarding the selected organisms, Genestrip produces k-mer databases and results comparable to those of KrakenUniq but at a fraction of its required memory resources. Our tool ensures that during database generation, the most suitable lowest common ancestor taxon is assigned for each stored k-mer by also considering genomes of organisms whose k-mers are not included in the database. This enables read analysis with high precision and recall for the organisms of interest.
Results:
We assess the correctness, usefulness and performance of Genestrip in different contexts and show that it indeed ascertains high quality read classifications for organisms whose genomes are included in a corresponding database. Our example databases comprise millions to a few billions of k-mers covering a dozen to a few thousands of species and lend themselves to usage in tick surveillance, medical diagnostics or agriculture. All databases were generated on a regular PC within hours, and related analysis performance was competitive to highly favorable. The deliberate focus on a particular set of genera or species allows for more genomes to be included from related organisms while the resulting databases remain small. Since k-mer compression becomes unnecessary, false positives emerging from related information loss are entirely avoided. We exemplify that such small but deep databases tend to improve recall during read classification while sustaining high precision.
Conclusions:
Due to Genestrip's particular way of updating the k-mers' lowest common ancestor taxa, both database creation and fastq file analysis can be realized with little memory and with favorable runtimes as well as high classification quality. So both, database creation and read classification may be performed even on regular PCs. Genestrip's qualities empower users to flexibly design, build and use small k-mer databases for their own needs with potentially deep genomic coverage.
More Related Videos
10:23A Concoction Pipeline for Generating Molecular Operational Taxonomic Units (MOTUs) Among Riparian and Aquatic Beetles
Published on: July 11, 2025
12:11Identification of Metabolically Active Bacteria in the Gut of the Generalist Spodoptera littoralis via DNA Stable Isotope Probing Using 13C-Glucose
Published on: November 13, 2013
Related Concept Videos
Methods of Classification and Identification
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Rapid Identification of Pathogens