POSA-GO:分层基因本体学和蛋白质语言模型的融合,用于蛋白质功能预测
Yubao Liu1, Benrui Wang1, Bocheng Yan1
1College of Computer Science and Technology, Changchun University, Changchun 130012, China.
International journal of molecular sciences
|July 12, 2025
概括
一种新的方法,基因本体学的部分依据自我注意 (POSA-GO),改善了蛋白质功能预测. 它有效地整合了蛋白质序列和基因本体学术语的等级结构,以获得更准确的功能注释.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 蛋白质功能预测对于理解后基因组时代的生物过程至关重要.
- 高通量测序比功能注释更快地生成数据,从而产生了对先进预测工具的关键需求.
- 现有的方法往往无法捕捉基因本体学 (GO) 术语和动态蛋白质特征关联的等级性.
研究的目的:
- 开发一种新的蛋白质功能预测框架,解决当前方法的局限性.
- 有效地建模GO术语的层次结构,并动态地将蛋白质特征与功能上下文联系起来.
- 为了提高蛋白质功能注释的准确性和效率.
主要方法:
- 提出了一个新的框架:基因本体学的部分基于顺序的自我注意力 (POSA-GO).
- 采用跨模式协作建模方法,将GO术语与蛋白质序列融合在一起.
- 利用预训练语言模型ESM-2进行蛋白序列特征提取,并将GO术语部分顺序关系转换为拓嵌入.
- 实施了多头自我注意机制,用于蛋白质和GO术语之间的动态关联重量建模.
主要成果:
- 与现有的最先进的方法相比,POSA-GO表现出更高的性能.
- 在基准数据集 (CAFA3和SwissProt) 上实现了改进的指标,特别是Fmax和AUPR.
- 该模型成功地捕获了GO术语之间的等级依赖关系,并使上下文感知功能注释成为可能.
结论:
- POSA-GO为准确的蛋白质功能预测提供了一个有希望和有效的解决方案.
- 该框架利用层次 GO 结构和动态特征关联的能力增强了功能注释.
- 这种方法推动了生物信息学领域的发展,并有助于破译分子机制.
相关概念视频
Protein Families
15.8K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.8K
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Protein-protein Interfaces
13.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
13.4K
Conservation of Protein Domains Over Different Proteins
11.4K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.4K
Protein and Protein Structure
81.5K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
81.5K
Tagging and Fusion Proteins
6.9K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.9K


