安静化器:一种灵活的神经网络框架,用于结构性注释和长读转录组的解复
Ayush Semwal1, Jacob Morrison1, Ian Beddows1
1Department of Epigenetics, Van Andel Research Institute, Grand Rapids, MI, USA.
bioRxiv : the preprint server for biology
|August 6, 2025
概括
在长时间读取的单细胞RNA测序数据中,Tranquillyzer准确地识别细胞条形码和独特的分子标识符. 这种深度学习框架处理序列错误和库变异,以便可靠地量化转录.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 长读单细胞RNA测序 (scRNA-seq) 提供了全长的转录组分析.
- 在长读scRNA-seq挑战细胞条形码 (CBC) 和独特分子标识符 (UMI) 识别中,错误率高和复杂的库结构.
- 现有的方法难以处理数据变化,如移位,截断和非正规排序.
研究的目的:
- 开发一个强大而灵活的深度学习框架来处理长时间读取的scRNA-seq数据.
- 为了准确地识别CBC和UMI,尽管有序列错误和库的变化.
- 提供一个可扩展的解决方案,用于分析大规模的长时间读取的转录数据集.
主要方法:
- 推出了Tranquillyzer,这是一个使用混合神经网络架构的深度学习框架.
- 采用全局,上下文意识的设计,精确识别结构元素.
- 启用了快速,一次性模型训练,用于自定义库格式.
主要成果:
- 镇定器准确地识别CBC和UMI,即使有移动,降解或重复的元素.
- 该框架在支持既定和定制单细胞协议方面表现出灵活性.
- 模型训练是高效的,通常在标准GPU上在几个小时内完成.
结论:
- 安静化器为长时间读取scRNA-seq数据处理提供了灵活,可扩展和准确的解决方案.
- 该框架克服了处理数据复杂性的现有方法的局限性.
- 它为增强的转录组分析提供了可靠的解复和脱复制.
更多相关视频
04:58Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance
Published on: December 13, 2024
2.9K
09:58Mapping the Structure-Function Relationships of Disordered Oncogenic Transcription Factors Using Transcriptomic Analysis
Published on: June 27, 2020
2.8K
相关概念视频
RNA-seq
10.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.4K
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
