Related Experiment Video
Updated: Mar 21, 2026

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
Phylogeny-aware identification and correction of taxonomically mislabeled sequences
Alexey M Kozlov1, Jiajie Zhang2, Pelin Yilmaz3
1The Exelixis Lab, Scientific Computing Group, Heidelberg Institute for Theoretical Studies, Schloss-Wolfsbrunnenweg 35, 69118 Heidelberg, Germany Alexey.Kozlov@h-its.org.
Automated identification and correction of taxonomic mislabels in molecular sequence databases are now possible. SATIVA, a new phylogeny-aware method, accurately detects and corrects erroneous sequence annotations, improving data reliability.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Public molecular sequence databases often contain taxonomic mislabels due to unvalidated author annotations.
- These errors propagate, affecting downstream analyses and biasing metagenetic studies.
- Manual curation is labor-intensive, leading to low correction rates.
Purpose of the Study:
- To develop and validate an automated method for identifying and correcting taxonomically mislabeled sequences.
- To assess the prevalence of mislabels in major microbial 16S rRNA gene reference databases.
- To evaluate taxonomic classifications using a phylogeny-aware approach.
Main Methods:
- Utilized the Evolutionary Placement Algorithm (EPA) to detect phylogenetic signal inconsistencies.
- Developed SATIVA, a phylogeny-aware method employing statistical models of evolution.
- Applied SATIVA to simulated data and real-world microbial 16S rRNA gene databases.
Main Results:
- SATIVA achieved high accuracy in identifying (96.9% sensitivity, 91.7% precision) and correcting (94.9% sensitivity, 89.9% precision) mislabeled sequences.
- Detected 0.2%–2.5% mislabels across Greengenes, LTP, RDP, and SILVA databases.
- Performed in-depth taxonomic evaluation for Cyanobacteria.
Conclusions:
- SATIVA offers an accurate and efficient solution for detecting and correcting taxonomic mislabels in sequence databases.
- The prevalence of mislabels in widely used databases highlights the need for automated quality control.
- Phylogeny-aware methods are crucial for reliable taxonomic assignment in molecular data.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Applications of Molecular Taxonomy
Microbial Phylogeny
Phylogenetic Trees
Phylogenetic Trees

