Related Experiment Video
Updated: May 20, 2026

09:30
Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms
Published on: September 13, 2018
Estimation of sequencing error rates in short reads
Xin Victoria Wang1, Natalie Blades, Jie Ding
1Department of Biostatistics and Computational Biology, Dana-Farber Cancer Institute, Boston, MA 02215, USA.
BMC Bioinformatics
|August 1, 2012
Summary
This study introduces a novel method for estimating sequencing error rates in short reads without needing a reference genome. This approach enhances the reliability of next-generation sequencing data analysis.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Next-generation sequencing (NGS) generates vast amounts of short-read data.
- Ensuring the fidelity of NGS data is crucial for research accuracy.
- Simple, reliable methods for experimental data quality monitoring are needed.
Purpose of the Study:
- To develop a method for estimating error rates in short-read sequencing data.
- To provide a tool for monitoring sequencing quality at the experiment level.
- To offer an alternative to reference-genome-dependent error estimation.
Main Methods:
- Developed a method based on the linear relationship between read copy number and erroneous reads.
- Error rate estimation by read and by position.
- Implementation as an R package for accessibility.
Main Results:
- Demonstrated a fast, scalable, and accurate error rate estimation approach.
- Method does not require a reference genome.
- Outperformed reference-genome-based methods in accuracy.
- Successfully detected mutations in PhiX strain genome data.
- Validated with simulation studies and real data sets.
Conclusions:
- The method enables quality monitoring of sequencing pipelines per experiment.
- Eliminates the need for reference genomes in error rate estimation.
- Provides error rate estimates to improve downstream NGS data analyses and inferences.
Related Concept Videos
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Sanger Sequencing
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...

