Related Experiment Video
Updated: Jun 5, 2025

Tick Microbiome Characterization by Next-Generation 16S rRNA Amplicon Sequencing
Published on: August 25, 2018
Impact of database choice and confidence score on the performance of taxonomic classification using Kraken2
Yunlong Liu1, Morteza H Ghaffari2, Tao Ma1
1Key Laboratory of Feed Biotechnology of the Ministry of Agricultural and Rural Affairs, Institute of Feed Research, Chinese Academy of Agricultural Sciences, Beijing, 100081 China.
Abstract:
Accurate taxonomic classification is essential to understanding microbial diversity and function through metagenomic sequencing. However, this task is complicated by the vast variety of microbial genomes and the computational limitations of bioinformatics tools. The aim of this study was to evaluate the impact of reference database selection and confidence score (CS) settings on the performance of Kraken2, a widely used k-mer-based metagenomic classifier. In this study, we generated simulated metagenomic datasets to systematically evaluate how the choice of reference databases, from the compact Minikraken v1 to the expansive nt- and GTDB r202, and different CS (from 0 to 1.0) affect the key performance metrics of Kraken2. These metrics include classification rate, precision, recall, F1 score, and accuracy of true versus calculated bacterial abundance estimation. Our results show that higher CS, which increases the rigor of taxonomic classification by requiring greater k-mer agreement, generally decreases the classification rate. This effect is particularly pronounced for smaller databases such as Minikraken and Standard-16, where no reads could be classified when the CS was above 0.4. In contrast, for larger databases such as Standard, nt and GTDB r202, precision and F1 scores improved significantly with increasing CS, highlighting their robustness to stringent conditions. Recovery rates were mostly stable, indicating consistent detection of species under different CS settings. Crucially, the results show that a comprehensive reference database combined with a moderate CS (0.2 or 0.4) significantly improves classification accuracy and sensitivity. This finding underscores the need for careful selection of database and CS parameters tailored to specific scientific questions and available computational resources to optimize the results of metagenomic analyses.
Supplementary Information:
The online version contains supplementary material available at 10.1007/s42994-024-00178-0.
More Related Videos
11:09Use of MALDI-TOF Mass Spectrometry and a Custom Database to Characterize Bacteria Indigenous to a Unique Cave Environment Kartchner Caverns, AZ, USA
Published on: January 2, 2015
08:20Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
Published on: October 27, 2023
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Confidence Coefficient
Kruskal-Wallis Test
Microbial Classification System
Applications of Molecular Taxonomy
Classification of Systems-II