Related Experiment Video
Updated: Jun 7, 2025

11:26
Sequencing of mRNA from Whole Blood using Nanopore Sequencing
Published on: June 3, 2019
13.6K
A mapping-free natural language processing-based technique for sequence search in nanopore long-reads
Tomasz Strzoda1, Lourdes Cruz-Garcia2, Mustafa Najim2
1Department of Data Science and Engineering, Silesian University of Technology, Gliwice, Poland.
BMC Bioinformatics
|November 13, 2024
Summary
A new natural language processing tool offers fast and accurate gene expression mapping for radiation dose estimation. This method shows high negative predictive value (NPV), outperforming classical approaches for long nanopore reads.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Accurate gene expression analysis is crucial for dose estimation in radiation emergencies.
- Existing methods for targeted gene mapping can be computationally intensive and energy-consuming.
- There is a need for rapid, low-energy, in-situ sequencing analysis tools.
Purpose of the Study:
- To develop a sequence identification tool using natural language processing (NLP) techniques.
- To achieve a high negative predictive value (NPV) for gene expression analysis.
- To create a computationally inexpensive and energy-efficient alternative to classical mapping methods.
Main Methods:
- Exploration of various NLP models with different dictionary components and encoding lengths.
- Training and validation of NLP models on RNA sequencing data.
- Comparison of NLP model performance against classical mapping tools like minimap2.
Main Results:
- The optimal NLP configuration analyzed entire sequences with a 3-base pair word length.
- The best model achieved 98.29% balanced accuracy (BACC) and 99.25% NPV for the FDXR gene.
- Reduced dictionary size maintained high NPV (98.15% internally, 99.64% externally), with minimal impact on read count accuracy.
Conclusions:
- NLP-based mapping is a reliable replacement for classical methods in targeted transcript analysis of long nanopore reads.
- The developed model is adaptable for different transcript targets and long-read sequencing technologies.
- Classical text processing techniques show significant potential for nucleotide sequence analysis.
Related Concept Videos
RNA-seq
9.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.8K
Next-generation Sequencing
87.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.6K
Sanger Sequencing
753.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
753.0K

