Slimformer:基于NLP的Web服务器,用于基因组的语义分类
Fionn Daire Keogh1, Jonas Marx2, Alicia Hiemisch1
1Institute of Infectious Diseases and Infection Control, Jena University Hospital, Jena, Germany.
Computational and structural biotechnology journal
|December 4, 2025
概括
Slimformer是一种新的自然语言处理工具,通过使用语义相似性对基因组进行分类来改进omics数据的解释. 它增强了对细胞机制的理解,正如呼吸道同胞细胞病毒研究所示.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 奥米克数据分析产生了大量的基因组,使细胞机制的解释变得复杂.
- 当前的分类方法往往忽视了文本描述中的语义相似性,严重依赖像基因本体学这样的等级本体学.
- 需要先进的方法来整合语言和功能信息来进行全面的基因组分析.
研究的目的:
- 开发和验证基于嵌入的自然语言处理模型Slimformer,用于增强基因组分类.
- 利用基因组名称,描述和相关基因的上下文关系来改进分类.
- 为系统的基因组分类提供灵活的框架,有助于对OMIC数据的解释.
主要方法:
- 开发了基于嵌入的自然语言处理模型Slimformer.
- 在手工策划的标注基因组的黄金标准数据集上训练了一个受监督的分类器.
- 利用基因组名,描述和相关基因来学习上下文关系.
- 将模型应用于2856个注释的基因组和RSV感染的人类细胞的基因表达数据.
主要成果:
- 在注释的基因组上,Slimformer实现了82.4%的平衡精度和0.867的F1得分.
- 该模型确定了RSV感染细胞中细胞循环过程的显著下调,这是其他工具错过的发现.
- 通过集成的语言和功能信息,证明了omics数据的可解释性.
结论:
- Slimformer通过结合语义相似性,为基因组分类提供了一种新且有效的方法.
- 该工具增强了omics数据的可解释性,促进了生物见解的发现.
- Slimformer为研究人员分析复杂的生物数据集和理解疾病机制提供了有价值的框架.
更多相关视频
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.9K
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
10.9K
相关概念视频
Gene Families
9.7K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
9.7K
Gene Families
3.5K
3.5K
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
