Related Experiment Video
Updated: Feb 20, 2026

12:08
Hybrid De Novo Genome Assembly for the Generation of Complete Genomes of Urinary Bacteria using Short- and Long-read Sequencing Technologies
Published on: August 20, 2021
5.9K
MegaGTA: a sensitive and accurate metagenomic gene-targeted assembler using iterative de Bruijn graphs
Dinghua Li1, Yukun Huang1, Chi-Ming Leung1,2
1Department of Computer Science, University of Hong Kong, Pokfulam, Hong Kong.
BMC Bioinformatics
|October 27, 2017
Summary
MegaGTA is a new gene-targeted assembler that improves upon Xander by using iterative de Bruijn graphs and succinct de Bruijn graphs (SdBG) for higher sensitivity, accuracy, and speed in metagenomic data analysis.
Area of Science:
- Metagenomics
- Bioinformatics
- Computational Biology
Background:
- Xander, a gene-targeted metagenomics assembler, utilizes Hidden Markov Models (HMMs) and de Bruijn graphs but has limitations in speed, k-mer size flexibility, and handling of false positives from Bloom filters.
- Xander's reliance on a single k-mer size compromises either sensitivity or accuracy, and its use of Bloom filters introduces potential assembly quality issues due to false positives.
Purpose of the Study:
- To develop a novel gene-targeted assembler, MegaGTA, that addresses the limitations of existing methods like Xander.
- To enhance the sensitivity, accuracy, and computational speed of metagenomic gene assembly.
Main Methods:
- MegaGTA employs iterative de Bruijn graphs to leverage multiple k-mer sizes, optimizing both sensitivity and accuracy.
- Utilizes succinct de Bruijn graphs (SdBG) for an exact representation of the de Bruijn graph, reducing memory footprint and enabling faster computation through parallel algorithms.
- Incorporates k-mer multiplicity for improved Hidden Markov Model (HMM) construction, avoiding false-positive contigs.
Main Results:
- MegaGTA demonstrated superior sensitivity and accuracy compared to Xander on mock metagenomic datasets.
- Analysis of a large rhizosphere soil metagenomic sample showed MegaGTA produced 9.7–19.3% more contigs and assigned 10–25% more gene references than Xander.
- MegaGTA exhibited a 2- to 10-fold increase in speed over Xander, depending on the number of k-mers used.
Conclusions:
- MegaGTA significantly improves upon Xander's gene assembly algorithms, delivering enhanced sensitivity, accuracy, and speed.
- The assembler is capable of efficiently assembling gene sequences from ultra-large metagenomic datasets.
- MegaGTA's source code is publicly available for further research and application.
Related Concept Videos
Genome Annotation and Assembly
21.1K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
21.1K
Next-generation Sequencing
99.2K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
99.2K
Sanger Sequencing
775.6K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
775.6K
Maxam-Gilbert Sequencing
13.1K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
13.1K
Genomics
41.0K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
41.0K
Gene Duplication and Divergence
8.1K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
8.1K

