Related Experiment Video
Updated: Jan 11, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.5K
ReadSeeker: A DNABERT based de-novo read-level gene predictor
Ben Wulf1, Piotr Wojciech Dabrowski1
1Center for Bio-Medical Image and Information Processing (CBMI), HTW University of Applied Sciences, Berlin, Berlin, Germany.
Plos One
|November 13, 2025
Summary
ReadSeeker, a new DNA model, accurately classifies DNA sequencing reads as protein-coding or non-protein-coding without reference genomes. This advancement aids in understanding genomic functions across various species.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Next-generation sequencing (NGS) generates vast amounts of short reads.
- Differentiating protein-coding (CDS) from non-protein-coding (non-CDS) regions is crucial for genomic annotation.
- Current methods often rely on reference sequences, limiting application in novel or unannotated genomes.
Purpose of the Study:
- To develop and evaluate ReadSeeker, a DNABERT-based model for classifying NGS short reads.
- To enable accurate CDS/non-CDS differentiation without reliance on reference genomes.
- To assess the model's performance across diverse biological datasets.
Main Methods:
- Fine-tuning a DNABERT model named ReadSeeker.
- Training on approximately 3 million synthetic reads from annotated viral, bacterial, and mammalian genomic elements.
- Evaluating performance on real-world human, viral, and bacterial sequencing data.
Main Results:
- ReadSeeker achieved high accuracy exceeding 94% in differentiating NGS short reads.
- Receiver Operating Characteristic Area Under the Curve (ROC-AUC) scores surpassed 98% in most evaluations.
- The model demonstrated robust performance across diverse sample types (human, viral, bacterial).
Conclusions:
- ReadSeeker provides a robust, reference-free method for classifying DNA sequencing reads.
- The model significantly advances genomic annotation capabilities, particularly for uncharacterized or novel sequences.
- High accuracy and AUC scores indicate ReadSeeker's potential for broad application in genomic research.
Related Concept Videos
Nonsense-mediated mRNA Decay
11.7K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
11.7K
RNA-seq
11.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.7K
Next-generation Sequencing
97.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.6K
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K

