Related Experiment Video
Updated: May 16, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
Musket: a multistage k-mer spectrum-based error corrector for Illumina sequence data
Yongchao Liu1, Jan Schröder, Bertil Schmidt
1Institut für Informatik, Johannes Gutenberg Universität Mainz, Mainz 55099, Germany. liuy@uni-mainz.de
Bioinformatics (Oxford, England)
|December 4, 2012
Summary
Musket is a new k-mer-based tool that efficiently corrects errors in Illumina short-read sequencing data. It achieves high precision and recall, improving genome assembly and demonstrating superior scalability.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Next-generation sequencing (NGS) technologies generate imperfect short-read data.
- Substitution errors are the primary error type in Illumina sequencing data.
- Existing short-read error correctors often achieve high recall or precision, but not both.
Purpose of the Study:
- To develop an efficient and accurate short-read error correction tool for Illumina data.
- To improve the quality of de novo genome assembly through enhanced read accuracy.
- To create a scalable and fast error correction method.
Main Methods:
- Developed Musket, a multistage k-mer-based error corrector.
- Employed a k-mer spectrum approach with three correction stages: two-sided conservative, one-sided aggressive, and voting-based refinement.
- Implemented a multi-threaded master-slave model for parallel processing.
Main Results:
- Musket demonstrates consistently top-tier performance in correction quality.
- The tool significantly improves de novo genome assembly metrics.
- Musket exhibits superior parallel scalability and competitive execution times compared to other correctors.
Conclusions:
- Musket is an efficient and highly scalable tool for correcting Illumina short-read sequencing errors.
- The developed correction techniques lead to improved genome assembly outcomes.
- Musket offers a robust solution for handling imperfect NGS data.
Related Concept Videos
Sanger Sequencing
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
NMR Spectrometers: Resolution and Error Correction
When magnetic nuclei in a sample achieve resonance and undergo relaxation, the signal detected in NMR is an approximately exponential free induction decay. Fourier transform of an exponential decay yields a Lorentzian peak in the frequency domain. Lorentzian peaks in an NMR spectrum are defined by their amplitude, full width at half maximum, and position, where the peak width is governed by the spin-spin relaxation time alone. In real experiments, however, the applied magnetic field is rendered...
Mismatch Repair
Overview
Mismatch Repair
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Mismatch Repair
Overview
Genome Copying Errors
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
