Related Experiment Video
Updated: Mar 25, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.7K
Iterative error correction of long sequencing reads maximizes accuracy and improves contig assembly
Briefings in Bioinformatics
|February 13, 2016
Summary
Correcting sequencing errors in long DNA reads is hard, especially in repeats. An iterative approach using String Graph Assembler (SGA) improves genome assembly accuracy and contig length.
Area of Science:
- Genomics
- Bioinformatics
Background:
- Next-generation sequencing (NGS) technologies produce long reads (up to 300 bp) beneficial for genome assembly.
- Computational error correction is a crucial first step in genome assembly but remains challenging for longer reads.
Purpose of the Study:
- To develop and evaluate an iterative error correction pipeline for long sequencing reads.
- To improve the accuracy of genome assembly by addressing errors within repetitive regions.
Main Methods:
- An iterative error correction pipeline was developed using modules from the String Graph Assembler (SGA).
- The pipeline employs multiple rounds of k-mer-based correction with increasing k-mer sizes, followed by overlap-based correction.
- This approach combines small and large k-mer advantages to correct errors in repeats.
Main Results:
- The iterative pipeline effectively corrects more errors within repeats compared to standard methods.
- Minimizing erroneous reads leads to a significant increase in contig lengths, two to three times longer.
- Higher read accuracy directly translates to improved genome assembly contiguity.
Conclusions:
- Iterative error correction is a robust strategy for enhancing the quality of long sequencing reads.
- The developed pipeline, SGA-Iteratively Correcting Errors, improves genome assembly by tackling errors in repetitive sequences.
- This method offers a valuable tool for researchers aiming for more contiguous and accurate genome assemblies.
Related Concept Videos
Genome Annotation and Assembly
21.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
21.5K
Proofreading
9.8K
Synthesis of new DNA molecules is carried out by the enzyme DNA polymerase, which adds nucleotides on the daughter strand complementary to the template DNA strand. DNA polymerase has a higher affinity to add the correct base and ensures fidelity during DNA replication. Furthermore, it exhibits proofreading activity during replication, using an exonuclease domain that cuts off incorrect nucleotides from the nascent DNA strand.
Errors During Replication are Corrected by the DNA Polymerase...
Errors During Replication are Corrected by the DNA Polymerase...
9.8K
Proofreading
62.0K
Overview
62.0K
Mismatch Repair
44.9K
Overview
44.9K
Mismatch Repair
7.0K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
7.0K
Improving Translational Accuracy
15.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.4K

