Related Experiment Video
Updated: May 27, 2026

09:30
Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms
Published on: September 13, 2018
RecountDB: a database of mapped and count corrected transcribed sequences.
Edward Wijaya1, Martin C Frith, Kiyoshi Asai
1Graduate School of Frontier Sciences, University of Tokyo, 5-1-5 Kashiwanoha, Kashiwa 277-8562, Japan.
Nucleic Acids Research
|December 6, 2011
Summary
RecountDB is a new database that corrects sequencing errors in gene expression data. This improves the detection of rare transcripts for more accurate transcriptomic analysis.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Next-generation sequencing (NGS) provides high-resolution gene expression data.
- NGS data contain errors that hinder accurate detection of rare transcripts.
Purpose of the Study:
- To present RecountDB, a secondary database for corrected gene expression data.
- To improve the accuracy of transcriptomic analysis by mitigating sequencer errors.
Main Methods:
- Derived secondary database from NCBI's Short Read Archive.
- Corrected and mapped RNA-sequencing (RNA-seq) and 5' capped transcription start site (TSS) data.
- Developed a searchable and browseable interface for data access.
Main Results:
- RecountDB contains corrected sequence counts from RNA-seq and TSS experiments.
- Data is mapped to the relevant genome and available in analysis-friendly formats.
- The database includes 2265 entries from 45 organisms and is expanding.
Conclusions:
- RecountDB offers corrected, high-quality gene expression data for transcriptomic studies.
- The database enhances the ability to detect and quantify rare transcripts.
- RecountDB is publicly accessible for researchers worldwide.
Related Concept Videos
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Cis-regulatory Sequences
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
Homologous Recombination
The basic reaction of homologous recombination (HR) involves two chromatids that contain DNA sequences sharing a significant stretch of identity. One of these sequences uses a strand from another as a template to synthesize DNA in an enzyme-catalyzed reaction. The final product is a novel amalgamation of the two substrates. To ensure an accurate recombination of sequences, HR is restricted to the S and G2 phases of the cell cycle. At these stages, the DNA has been replicated already and the...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Polytene Chromosomes
Polytene chromosomes are giant interphase chromosomes with several DNA strands placed side by side. They were discovered in the year 1881 by Balbiani in salivary glands, intestine, muscles, malpighian tubules, and hypoderm of larvae Chironomus plumosus. Hence, these are also called "Salivary gland chromosomes." These are found in insects of the order Diptera and Collembola; in certain organs of mammals; and synergids, antipodes of flowering plants. Polytene chromosomes are also regularly...
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.

