FGeneBERT:功能驱动的预训练基因语言模型用于元基因组学
Chenrui Duan1,2, Zelin Zang3, Yongjie Xu1,2
1College of Computer Science and Technology, Zhejiang University, No. 866, Yuhangtang Road, 310058 Zhejiang, P. R. China.
Briefings in bioinformatics
|November 10, 2025
概括
FGeneBERT是一种新的元基因组模型,使用基于蛋白质的基因上下文来改善对基因关系和功能的理解. 这种方法在复杂的生物数据中增强了基因,功能,细菌和环境层面的分析.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 超基因组数据对于了解多样化的环境和人类健康至关重要,但由于混合基因组,这也带来了挑战.
- 目前基于K-mer的方法限制了捕捉结构和功能基因背景,并与复杂的基因关系作斗争.
- 现有的方法无法有效地编码具有生物意义的基因,并解决元基因组数据中固有的一对多/多对一关系.
研究的目的:
- 介绍FGeneBERT,这是一种新的预训练模型,旨在克服当前元基因组数据分析的局限性.
- 增强对基因间上下文关系和基因序列功能关联的理解.
- 改进复杂的元基因组数据集中的生物意义基因的表示和分析.
主要方法:
- 开发了FGeneBERT,这是一种利用基于蛋白质的基因表征作为标记器的元基因学预训练模型.
- 实施了掩盖基因建模,以提高对基因间情境关系的理解.
- 采用三重增强的元基因学对比学习来阐明基因序列功能关系.
主要成果:
- 在基因,功能,细菌和环境层面上,FGeneBERT在超基因组数据集中表现出卓越的性能.
- 该模型有效地处理不同的输入序列大小,从1k到213k.
- 关于ATP合成酶和基因操作子的案例研究展示了FGeneBERT在功能识别和生物相关性方面的准确性.
结论:
- FGeneBERT通过提供上下文感知和结构相关的基因表示,在元基因组数据分析方面取得了重大进展.
- 基于蛋白质的方法和高级学习策略增强了元基因组发现的生物解释性.
- FGeneBERT的能力对于未来的环境和人类健康相关的元基因组学研究至关重要.
相关概念视频
Genetic Lingo
113.6K
Overview
113.6K
Genomics
39.6K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
39.6K
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Gene Families
3.5K
3.5K


