Related Experiment Video
Updated: Dec 23, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.4K
Overlap detection on long, error-prone sequencing reads via smooth q-gram
Yan Song1, Haixu Tang1, Haoyu Zhang1
1Department of Computer Science, Indiana University Bloomington, Bloomington, IN 47405, USA.
Bioinformatics (Oxford, England)
|April 21, 2020
Summary
A new smooth q-gram algorithm improves overlap detection for long, error-prone sequencing reads from PacBio and Nanopore technologies, especially for shorter overlaps.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Third-generation sequencing (PacBio, Nanopore) produces long, error-prone reads.
- Accurate overlap detection is crucial for de novo fragment assembly.
- Current methods struggle with high error rates and short overlaps.
Purpose of the Study:
- To address limitations in overlap detection for error-prone sequencing reads.
- To develop a novel algorithm for improved overlap detection accuracy.
Main Methods:
- Proposed 'smooth q-gram', a q-gram variant capturing pairs within small edit distances.
- Designed a novel algorithm utilizing smooth q-gram-based seeds for overlap detection.
- Implemented and tested the algorithm on PacBio and Nanopore datasets.
Main Results:
- The smooth q-gram algorithm outperforms existing q-gram-based methods.
- Significant improvements observed particularly for reads with short overlapping lengths.
- Demonstrated enhanced precision and recall in overlap detection.
Conclusions:
- The novel smooth q-gram approach offers superior performance for overlap detection.
- This method is particularly beneficial for analyzing challenging, error-prone sequencing data.
- The developed algorithm advances fragment assembly for third-generation sequencing.
Related Concept Videos
Detection of Gross Error: The Q Test
6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K
Sanger Sequencing
772.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.0K
RNA-seq
11.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.6K
Genome Copying Errors
4.9K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.9K
Improving Translational Accuracy
13.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
13.9K
Improving Translational Accuracy
3.4K
3.4K

