Related Experiment Video
Updated: Feb 8, 2026

The ITS2 Database
Published on: March 12, 2012
Removing contaminants from databases of draft genomes.
Jennifer Lu1,2, Steven L Salzberg1,2,3
1Department of Biomedical Engineering, Johns Hopkins University, Baltimore, MD, United States of America.
Metagenomic sequencing aids infection diagnosis but suffers from genome contamination. This study developed a bioinformatics system to create a clean eukaryotic pathogen genome database, improving diagnostic accuracy and reducing false positives.
Area of Science:
- Microbiology
- Bioinformatics
- Genomics
Background:
- Metagenomic sequencing of patient samples is a powerful tool for diagnosing human infections by capturing all DNA/RNA from pathogens.
- Accurate pathogen identification relies on high-quality reference genomes, but contamination in public databases leads to false positives.
- Existing genomic databases for eukaryotic pathogens contain contaminants (human, bacterial, archaeal, viral) and low-complexity sequences, compromising diagnostic reliability.
Purpose of the Study:
- To develop and validate a bioinformatics system for removing contamination and low-complexity sequences from eukaryotic pathogen genomes.
- To create a curated, "clean" database of eukaryotic pathogen genomes for improved metagenomic analysis.
- To enhance the accuracy and reliability of pathogen detection in clinical metagenomic samples.
Main Methods:
- Developed a bioinformatics pipeline to identify and remove non-target DNA/RNA sequences (human, bacterial, archaeal, viral) from draft eukaryotic pathogen genomes.
- Implemented filtering for low-complexity genomic sequences to eliminate another source of false positives.
- Applied the pipeline to a comprehensive database of sequenced eukaryotic pathogen genomes to generate a cleaned dataset.
Main Results:
- Successfully produced a database of "clean" eukaryotic pathogen genomes by removing contaminating sequences and low-complexity regions.
- Demonstrated that the new database significantly reduces false positives when identifying eukaryotic pathogens in metagenomic samples.
- Showcased improved sensitivity in detecting eukaryotic pathogens compared to using original, uncleaned genome databases.
Conclusions:
- The developed bioinformatics system effectively purges contamination and low-complexity sequences from eukaryotic pathogen genome databases.
- The resulting "clean" genome database enhances the accuracy and sensitivity of metagenomic diagnostics for human infections.
- This work provides a critical resource for improving the reliability of pathogen identification in clinical metagenomic sequencing.
More Related Videos
07:50Detection and Removal of Nuclease Contamination During Purification of Recombinant Prototype Foamy Virus Integrase
Published on: December 8, 2017
08:49Vegetated Treatment Systems for Removing Contaminants Associated with Surface Water Toxicity in Agriculture and Urban Runoff
Published on: May 15, 2017
Related Concept Videos
Genomics
Contaminants and Errors
Another key consideration is determining the appropriate number of samples required to...
Genomic Imprinting and Inheritance
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
Genome Size and the Evolution of New Genes
Comparing Mitochondrial, Chloroplast, and Prokaryotic Genomes
Extracorporeal Removal of Drugs: Hemoperfusion and Hemofiltration