相关实验视频
Updated: Jan 11, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.5K
阅读寻找者:一个基于DNABERT的de-novo阅读水平基因预测器
Ben Wulf1, Piotr Wojciech Dabrowski1
1Center for Bio-Medical Image and Information Processing (CBMI), HTW University of Applied Sciences, Berlin, Berlin, Germany.
PloS one
|November 13, 2025
概括
新的DNA模型ReadSeeker准确地将DNA测序读取分类为蛋白质编码或非蛋白质编码,没有参考基因组. 这一进步有助于理解各种物种的基因组功能.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 下一代测序 (NGS) 产生了大量的短读数.
- 区分蛋白质编码 (CDS) 和非蛋白质编码 (非CDS) 区域对于基因组注释至关重要.
- 目前的方法通常依赖于参考序列,限制了在新型或未注释的基因组中的应用.
研究的目的:
- 开发和评估基于DNABERT的ReadSeeker模型,用于NGS短读的分类.
- 在不依赖参考基因组的情况下实现准确的CDS/非CDS差异化.
- 评估模型在各种生物数据集中的表现.
主要方法:
- 微调一个名为ReadSeeker的DNABERT模型.
- 培训大约300万个合成阅读来自注释的病毒,细菌和哺乳动物基因组元素.
- 评估真实世界的人类,病毒和细菌测序数据的性能.
主要成果:
- 在区分NGS短阅读时,ReadSeeker实现了超过94%的高精度.
- 在大多数评估中,接收器运行特征曲线下面的区域 (ROC-AUC) 得分超过了98%.
- 该模型在各种样本类型 (人类,病毒,细菌) 中表现出强的性能.
结论:
- ReadSeeker提供了一种强大的,无引用的方法来分类DNA测序阅读.
- 该模型显著提升了基因组注释能力,特别是在未表征或新型序列方面.
- 高精度和AUC分数表明ReadSeeker在基因组研究中具有广泛应用的潜力.
相关概念视频
Nonsense-mediated mRNA Decay
11.7K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
11.7K
RNA-seq
11.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.7K
Next-generation Sequencing
97.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.6K
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K

