整合基因本体学关系用于使用PFresGO预测蛋白质功能
Tong Pan1, Geoffrey I Webb2, Seiya Imoto3,4
1Biomedicine Discovery Institute and Department of Biochemistry and Molecular Biology, Monash University, Melbourne, VIC, Australia.
Methods in molecular biology (Clifton, N.J.)
|July 29, 2025
概括
PFresGO是一个新的深度学习工具,通过利用基因本体学图的等级结构来预测多个蛋白质功能. 这种方法通过考虑函数之间的关系来改进现有方法,以便准确的高通量注释.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 高通量测序产生了大量的数据,但确定蛋白质功能仍然是一个挑战.
- 现有的计算方法通常会单独预测函数,忽视它们之间的复杂关系.
- 了解蛋白质功能对于生物研究和药物发现至关重要.
研究的目的:
- 引入PFresGO,一种基于注意力的深度学习模型,用于预测多个蛋白质功能.
- 利用基因本体学 (GO) 图的等级结构,以改善功能注释.
- 为蛋白质功能预测提供高通量和准确的方法.
主要方法:
- 开发了一个以注意力为基础的深度学习架构,命名为PFresGO.GO.
- 将基因本体学 (GO) 图的等级结构纳入模型.
- 应用该模型从序列数据同时预测多个蛋白质功能.
主要成果:
- PFresGO以高通量的方式准确地预测多种蛋白质功能.
- 该模型有效地利用GO图表中的等级关系,以提高预测准确度.
- 证明了预测结果的可解释性.
结论:
- PFresGO为蛋白质功能注释提供了一种高效准确的解决方案.
- 整合GO图表层次结构显著改善了多标签函数预测.
- 对于处理大规模蛋白质序列数据的研究人员来说,PFresGO是一个有价值的工具.
相关概念视频
Protein-protein Interfaces
13.3K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
13.3K
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Protein Families
3.4K
3.4K
Conservation of Protein Domains Over Different Proteins
11.4K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.4K
Tagging and Fusion Proteins
6.9K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.9K
Protein Folding Quality Check in the RER
3.8K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.8K


