Related Experiment Video
Updated: Jun 25, 2025

05:45
Validating Whole Genome Nanopore Sequencing, using Usutu Virus as an Example
Published on: March 11, 2020
8.8K
Sequencing accuracy and systematic errors of nanopore direct RNA sequencing.
Wang Liu-Wei1,2,3, Wiep van der Toorn4,5, Patrick Bohn6
1Systems Medicine of Infectious Disease (P5), Robert Koch Institute, Berlin, Germany. liuwei.wang@fu-berlin.de.
BMC Genomics
|May 28, 2024
Summary
Direct RNA sequencing (dRNA-seq) offers full-length transcripts but has understudied accuracy. This study reveals consistent error patterns, with deletions common and cytosine/uracil-rich regions prone to errors, impacting RNA sequencing data quality.
Area of Science:
- Genomics
- Molecular Biology
- Bioinformatics
Background:
- Direct RNA sequencing (dRNA-seq) on Oxford Nanopore Technologies (ONT) platforms enables full-length transcript sequencing, including modifications and poly-A tail information.
- Despite its potential, the accuracy and error profiles of dRNA-seq remain insufficiently characterized.
Purpose of the Study:
- To comprehensively evaluate dRNA-seq accuracy and characterize systematic errors across diverse organisms and synthetic RNAs.
- To identify sequence contexts and signal-level features contributing to sequencing errors.
- To assess the impact of different basecallers and identify sources of data loss.
Main Methods:
- Analysis of dRNA-seq data from multiple species and synthetic RNAs using ONT kits SQK-RNA001 and SQK-RNA002.
- Systematic characterization of error types (deletions, insertions, mismatches) and their distribution.
- Examination of raw signal data to correlate errors with sequence context and signal features.
- Evaluation of basecaller performance and identification of adapter-related errors.
- Generation and analysis of dRNA-seq data using the SQK-RNA004 kit.
Main Results:
- Median read accuracy ranged from 87% to 92%, with deletions being the predominant error type.
- Heteropolymers and short homopolymers were major error contributors; cytosine/uracil-rich regions showed higher error rates.
- Systematic biases were observed at nucleotide and motif levels, linked to signal-level features.
- Read quality scores partially predict error rates; undetected DNA adapters contribute to errors.
- The newer SQK-RNA004 kit improved overall accuracy but retained similar systematic error patterns.
Conclusions:
- This study provides the first systematic analysis of dRNA-seq errors, revealing reproducible patterns.
- Identified signal-level insufficiencies and sequence context dependencies as sources of errors.
- Lays the groundwork for developing robust error correction methods for dRNA-seq data.
Related Concept Videos
RNA-seq
9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Next-generation Sequencing
88.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.7K
Sanger Sequencing
754.1K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.1K

