Related Experiment Video
Updated: Mar 19, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.7K
Sparc: a sparsity-based consensus algorithm for long erroneous sequencing reads.
1Department of Computer Science, University of Maryland , College Park, MD , USA.
Peerj
|June 23, 2016
Summary
Sparc is a new algorithm that creates accurate genome sequences from long but error-prone third-generation sequencing (3GS) data. It efficiently combines 3GS and next-generation sequencing (NGS) data for high-quality genome assembly.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Third-generation sequencing (3GS) offers long reads but has high error rates (15-40%).
- Next-generation sequencing (NGS) provides high accuracy but shorter reads.
- Accurate genome assembly and variant calling require high-quality sequences from 3GS data.
Purpose of the Study:
- To develop an efficient algorithm for generating high-quality consensus sequences from 3GS data.
- To facilitate de novo genome assembly using long but erroneous sequencing reads.
- To enable the combined use of 3GS and NGS data for improved genome analysis.
Main Methods:
- Developed Sparc, a linear complexity consensus algorithm.
- Constructs a sparse k-mer graph from sequencing reads.
- Identifies the consensus sequence by finding the heaviest path in a reweighted graph.
Main Results:
- Sparc achieves consensus error rates below 0.5% with 30× PacBio data.
- Combined 3GS (Oxford Nanopore) and NGS data yield similar high-quality results.
- Sparc uses 80% less memory and time compared to existing methods.
- Demonstrates significant improvements in cost and computational efficiency.
Conclusions:
- Sparc effectively generates high-quality consensus sequences from 3GS data.
- The algorithm efficiently integrates 3GS and NGS data for superior genome assembly.
- Sparc offers a computationally efficient and accurate solution for genomic analysis.
Related Concept Videos
Multi-species Conserved Sequences
4.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.9K
RNA-seq
12.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.4K
Maxam-Gilbert Sequencing
13.5K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
13.5K

