Related Experiment Video
Updated: Apr 17, 2026

12:08
Hybrid De Novo Genome Assembly for the Generation of Complete Genomes of Urinary Bacteria using Short- and Long-read Sequencing Technologies
Published on: August 20, 2021
6.1K
DIME: a novel framework for de novo metagenomic sequence assembly
Xuan Guo1, Ning Yu, Xiaojun Ding
11 Departments of Computer Science and Biology, Georgia State University , Atlanta, Georgia .
Summary
A new metagenomic assembly framework, DIME (DIvide, conquer, and MErge), improves accuracy and efficiency for large datasets. It reconstructs more bases and yields higher quality assemblies, even with low-coverage data.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Next-generation sequencing platforms have increased metagenomic data size and decreased analysis costs.
- Existing assemblers struggle with large datasets, particularly those with low coverage and many non-overlapping contigs, leading to a trade-off between contig noise and sequence length.
Purpose of the Study:
- To develop a novel metagenomic sequence assembly framework, DIME, addressing limitations in accuracy and efficiency for large-scale projects.
- To provide MapReduce implementations of DIME for parallel processing on the Apache Hadoop platform.
Main Methods:
- Developed the DIME (DIvide, conquer, and MErge) assembly framework.
- Implemented DIME using MapReduce on Apache Hadoop (DIME-cap3 and DIME-genovo).
- Compared DIME against five popular assemblers (Cap3, Genovo, MetaVelvet, SOAPdenovo, SPAdes) on synthetic and real metagenomic datasets.
Main Results:
- DIME accurately partitions sequence reads.
- DIME reconstructs more bases and generates higher quality consensus sequences.
- DIME achieved higher assembly scores (corrected N50, BLAST-score-per-base) compared to other tools.
- DIME demonstrated near-theoretical speed-up on the Hadoop platform.
Conclusions:
- DIME offers significant improvements in metagenomic assembly, enhancing both accuracy and efficiency.
- The framework is robust to decreasing sequence coverage and performs well across various sequence abundances.
- DIME is a promising tool for analyzing large-scale metagenomic datasets.
Related Concept Videos
Genome Annotation and Assembly
22.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
22.3K
Next-generation Sequencing
102.2K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
102.2K
RNA-seq
12.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.7K
Sanger Sequencing
780.9K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
780.9K
Genomics
42.0K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
42.0K
RACE - Rapid Amplification of cDNA Ends
7.7K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
7.7K

