Related Experiment Video
Updated: Jan 3, 2026

11:22
Microbiota Analysis Using Two-step PCR and Next-generation 16S rRNA Gene Sequencing
Published on: October 15, 2019
31.0K
Struo: a pipeline for building custom databases for common metagenome profilers
Jacobo de la Cuesta-Zuluaga1, Ruth E Ley1, Nicholas D Youngblut1
1Department of Microbiome Science. Max Planck Institute for Developmental Biology, Tübingen 72076, Germany.
Bioinformatics (Oxford, England)
|November 29, 2019
Summary
Struo is a new pipeline that automates the creation of custom microbial genome databases for metagenome profiling. This improves the accuracy of analyzing microbial communities by incorporating the latest genomic data.
Area of Science:
- Microbiology
- Bioinformatics
- Genomics
Background:
- Metagenome profiling relies on gene and genome databases for analyzing microbial communities.
- Current databases lag behind the rapid increase in available microbial genomes, hindering accurate analysis.
- Unifying database content across different metagenome profiling tools is challenging.
Purpose of the Study:
- To develop an automated pipeline for constructing custom metagenome databases.
- To enhance the accuracy and efficiency of microbial community analysis.
- To address the limitations of static, outdated databases in metagenome profiling.
Main Methods:
- Developed Struo, a modular pipeline for automated genome acquisition from public repositories.
- Enabled the construction of custom databases compatible with multiple metagenome profilers.
- Incorporated novel genomes to represent broader microbial diversity.
Main Results:
- Struo automates the process of creating up-to-date, comprehensive genome databases.
- Custom databases significantly increased the mappability of sequence reads in both synthetic and real metagenome datasets.
- Improved representation of known microbial diversity in databases led to substantial gains in data analysis.
Conclusions:
- Struo provides an efficient solution for generating custom metagenome databases.
- Utilizing up-to-date, diverse genomic data enhances the reliability of microbial community profiling.
- The pipeline facilitates more accurate taxonomic and functional insights from metagenomic data.

