Related Experiment Video
Updated: May 21, 2026

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
The M5nr: a novel non-redundant database containing protein sequences and annotations from multiple sources and
Andreas Wilke1, Travis Harrison, Jared Wilkening
1Mathematics and Computer Science Division, Argonne National Laboratory, Argonne, IL 60439, USA.
Metagenome analysis is limited by computation time for sequence similarity. We present a shared reference database and tools to translate similarity searches, enabling reusable computational results for large datasets.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Sequence similarity computation is a bottleneck in metagenome analysis.
- Sharing similarity results requires a common reference to avoid recomputation.
- Standardized formats for similarity data can reduce computational burden.
Purpose of the Study:
- To develop a mechanism for maintaining a comprehensive, non-redundant protein database.
- To create tools for translating similarity search results into various annotation namespaces.
- To enable sharing of computational results for large sequence datasets.
Main Methods:
- Automated maintenance of a comprehensive protein database with quarterly releases.
- Development of tools to translate similarity searches into annotation namespaces like KEGG and GenBank.
- Creation of a common reference for sequence similarity data.
Main Results:
- A continuously updated, non-redundant protein database.
- Tools for flexible translation of similarity search outputs.
- Facilitation of data sharing and reduced computational redundancy.
Conclusions:
- The presented data and tools allow for the generation of multiple result sets from a single computation.
- Enables computational results to be shared across different research groups for large sequence datasets.
- Addresses the computational limitations in metagenome analysis through shared resources and standardized outputs.
More Related Videos
07:38Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
Published on: April 11, 2019
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Genome Annotation and Assembly
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...