Metagenome Proteins and Database Contamination

Irina R Arkhipova1

  • 1Josephine Bay Paul Center for Comparative Molecular Biology and Evolution, Marine Biological Laboratory, Woods Hole, Massachusetts, USA iarkhipova@mbl.edu.

Msphere
|November 5, 2020
PubMed

Insights

Metagenomic data contains many misannotated proteins, which harms the value of taxonomic identifiers in databases like RefSeq. Both data submitters and database managers must act to ensure data accuracy.

Area of Science:

  • Genomics
  • Bioinformatics
  • Proteomics

Background:

  • Metagenomic data analysis relies heavily on accurate taxonomic identification.
  • Current protein databases, such as RefSeq, are increasingly incorporating metagenome-derived proteins.
  • Misannotation of taxonomy in these proteins poses a significant challenge to data integrity.

Purpose of the Study:

  • To highlight the problem of misannotated taxonomy in metagenome-derived proteins.
  • To emphasize the threat these errors pose to the utility of taxonomic identifiers.
  • To call for immediate action from data submitters and database managers.

Main Methods:

  • Analysis of protein databases for taxonomic annotation accuracy.
  • Review of data submission protocols for metagenomic datasets.
  • Assessment of database management strategies for error correction.

Main Results:

  • A significant influx of metagenome-derived proteins with incorrect taxonomic labels into public databases.
  • Compromised reliability of taxonomy identifiers due to widespread misannotations.
  • Identification of RefSeq as a database affected by this issue.

Conclusions:

  • Urgent interventions are required to address the misannotation of metagenomic protein data.
  • Collaborative efforts between metagenomic data submitters and database curators are essential.
  • Maintaining the integrity of taxonomic identifiers is crucial for the future of biological data analysis.

Related Concept Videos

Genome Annotation and Assembly03:36

Genome Annotation and Assembly

The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.0K
Genomic DNA in Prokaryotes00:46

Genomic DNA in Prokaryotes

The genome of most prokaryotic organisms consists of double-stranded DNA organized into one circular chromosome in a region of cytoplasm called the nucleoid. The chromosome is tightly wound, or supercoiled, for efficient storage. Prokaryotes also contain other circular pieces of DNA called plasmids. These plasmids are smaller than the chromosome and often carry genes that confer adaptive functions, such as antibiotic resistance.
Genomic Diversity in Bacteria
Although bacterial genomes are much...
47.4K
Genome Size and the Evolution of New Genes03:21

Genome Size and the Evolution of New Genes

3.0K
Genome Size and the Evolution of New Genes03:21

Genome Size and the Evolution of New Genes

While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
8.8K
Genome Copying Errors02:46

Genome Copying Errors

DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their  survival. Therefore, the copying errors are checked and repaired at three levels.
4.8K
Genomics02:02

Genomics

Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
39.0K