Related Experiment Video
Updated: Jun 12, 2026

10:41
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved (Non-model) Organisms
Published on: May 9, 2017
mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies
Tatiana A Gurbich1, Martin Beracochea1, Nishadi H De Silva1
1European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Cambridge CB10 1SD, United Kingdom.
Bioinformatics (Oxford, England)
|January 24, 2025
Summary
Mettannotator is a scalable Nextflow pipeline for annotating prokaryotic genomes, including novel species. It identifies genes, predicts functions like antimicrobial resistance, and outputs results for downstream analysis.
Area of Science:
- Genomics
- Bioinformatics
- Microbial Ecology
Background:
- Increasing numbers of prokaryotic genome assemblies from isolates and environmental samples present challenges for annotation.
- Novel species are often poorly represented in existing reference databases, necessitating advanced annotation tools.
Purpose of the Study:
- To introduce mettannotator, a comprehensive and scalable Nextflow pipeline for prokaryotic genome annotation.
- To enable annotation of both well-described and novel prokaryotic taxa at scale.
Main Methods:
- The mettannotator pipeline is implemented in Nextflow and Python.
- It identifies coding and noncoding regions, predicts protein functions (including antimicrobial resistance), and delineates gene clusters.
- Results are summarized in a General Feature Format (GFF) file.
Main Results:
- The pipeline was tested on 200 genomes from 29 prokaryotic phyla, including isolate and metagenome-assembled genomes.
- Performance metrics were generated and compared against other annotation tools.
- The tool successfully annotates diverse prokaryotic genomes, including novel taxa.
Conclusions:
- Mettannotator provides a robust and scalable solution for prokaryotic genome annotation.
- The pipeline facilitates downstream analysis and visualization of genomic data.
- It addresses the need for efficient annotation of both known and novel prokaryotic species.
Related Concept Videos
Next-generation Sequencing
87.3K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.3K
RNA-seq
9.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.8K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K

