Detecting and correcting misclassified sequences in the large-scale public databases.

Hamid Bagheri1, Andrew J Severin2, Hridesh Rajan1

  • 1Department of Computer Science, Ames, IA 50011, USA.

Summary

A new method identified over two million misclassified proteins in the non-redundant (NR) database. This approach offers high precision for detecting taxonomic errors in large biological sequence datasets.