Related Experiment Video
Updated: Jul 14, 2025

12:37
Efficient Nucleic Acid Extraction and 16S rRNA Gene Sequencing for Bacterial Community Characterization
Published on: April 14, 2016
38.7K
A streamlined pipeline based on HmmUFOtu for microbial community profiling using 16S rRNA amplicon sequencing
Hyeonwoo Kim1, Jiwon Kim1, Ji Won Choi2,3
1Department of Bioinformatics, Soongsil University, Seoul 06978, Korea.
Genomics & Informatics
|October 9, 2023
Summary
This study introduces a streamlined pipeline for microbial 16S rRNA amplicon sequencing, showing that HmmUFOtu retains significantly more reads than DADA2 for large datasets. This method also enables effective merging of microbial community datasets.
Area of Science:
- Microbiology
- Bioinformatics
- Genomics
Background:
- 16S rRNA amplicon sequencing is crucial for microbial community profiling.
- Amplicon Sequence Variant (ASV) methods offer high resolution but often discard many reads in large datasets.
- Operational Taxonomic Unit (OTU) clustering methods may offer better read retention.
Purpose of the Study:
- To develop and evaluate a streamlined bioinformatics pipeline for large-scale 16S rRNA amplicon sequencing data.
- To compare the read retention and performance of an OTU-based method (HmmUFOtu) against a common ASV method (DADA2).
- To assess the capability of the pipeline for merging datasets and identifying batch effects.
Main Methods:
- A pipeline integrating FastP, HmmUFOtu, Vsearch, and Kraken2 was developed.
- Two large published stool datasets (890 and 1,462 samples) from Korean populations were reprocessed.
- Performance was evaluated based on read retention, taxonomic assignment, and beta-diversity analysis.
Main Results:
- HmmUFOtu retained 93.2% and 89.2% of reads in the two datasets, significantly higher than DADA2's 44.6% and 18.4%.
- Both methods produced qualitatively similar beta-diversity patterns.
- HmmUFOtu facilitated dataset merging with high abundance correlation (R=0.92), though a batch effect was detected on the third dimension of beta-diversity.
Conclusions:
- The streamlined pipeline, particularly using HmmUFOtu, offers superior read retention for large-scale microbial 16S rRNA amplicon sequencing.
- HmmUFOtu's closed-reference approach simplifies merging of independent datasets.
- The findings provide insights into optimizing large dataset processing and highlight potential batch effects.

