通过学习和整合蛋白质序列和功能标签的表示来改善蛋白质功能预测.
Frimpong Boadu1, Jianlin Cheng1
1Department of Electrical Engineering and Computer Science, NextGen Precision Health Institute, University of Missouri, Columbia, MO 65211, United States.
预测蛋白质功能是一项挑战,特别是在罕见的术语. 新型变压器模型TransFew通过整合蛋白序列和基因本体学 (GO) 术语来提高蛋白质功能预测的准确性.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 对不到1%的蛋白质来说,实验性地确定蛋白质功能是可行的,对于大多数蛋白质来说,需要计算预测.
- 目前的蛋白质功能预测方法难以准确,特别是对于罕见的基因本体学 (GO) 术语,在数据库中注释有限,如UniProt.
研究的目的:
- 介绍TransFew,一种新型的变压器模型,旨在提高蛋白质功能预测的准确性.
- 通过利用蛋白质序列和GO术语的集成表示来改善罕见蛋白质功能术语的预测.
主要方法:
- 对于蛋白质序列表示,TransFew使用预训练的蛋白质语言模型 (ESM2-t48).
- 它采用生物自然语言模型 (BioBert) 和基于图形卷积神经网络的自编码器,用于GO术语的语义表示,考虑其定义和层次关系.
- 交叉注意力机制整合了蛋白质序列和GO术语表示,用于功能预测.
主要成果:
- 在TransFew的综合方法显著提高了整体蛋白质功能预测的准确性.
- TransFew在预测具有有限注释的罕见函数术语方面表现出强的表现.
- 该模型促进了GO术语之间的注释转移,改善了代表性不足的函数的预测.
结论:
- TransFew为准确的蛋白质功能预测提供了一种强大的新方法,解决现有方法的局限性.
- 该模型处理罕见函数术语的能力对于注释绝大多数具有未知功能的蛋白质具有重大意义.
- 序列和标签表示的集成是提高TransFew性能的关键.
更多相关视频
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
相关概念视频
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Families
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
