GOBoost:利用长尾基因本体学术语进行准确的蛋白质功能预测
Lei Zhang1, Yang Wang1, Xiao Chen2
1Department of Computer Science and Technology, Anhui University, Hefei, Anhui 230601, China.
Bioinformatics (Oxford, England)
|June 28, 2025
概括
GOBoost通过解决基因本体学术语的长尾分布来增强蛋白质功能预测. 这种深度学习方法的性能优于最先进的方法,提高了预测蛋白质功能的准确性.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 机器学习在生物学中的应用
背景情况:
- 对于蛋白质功能预测的深度学习方法往往忽视了基因本体学术语的长尾分布.
- 这种不平衡会导致低频函数的预测精度低于最佳.
研究的目的:
- 提出GOBoost,一种用于蛋白质功能预测的新方法,该方法明确地解决了功能标签的长尾分布.
- 提高蛋白质功能的预测准确度,特别是对于罕见或代表性不足的类别.
主要方法:
- GOBoost采用了一个长尾优化整体策略.
- 它集成了全球-本地标签图模块和多颗粒度焦点损失功能.
- 这些组件旨在更好地捕获和利用长尾功能信息.
主要成果:
- 在所有评估指标中,GOBoost在PDB和AF2数据集上显著超过了最先进的方法.
- 具体来说,GOBoost显示了AUPR对分子功能 (MF),生物过程 (BP) 和细胞组件 (CC) 预测的实质性改进.
- 结果强调了在模型设计中考虑标签分布的重要性.
结论:
- 与现有方法相比,GOBoost方法在蛋白质功能预测方面表现出优异的性能.
- 解决基因本体学术语的长尾分布对于提高计算函数预测模型的准确性和稳定性至关重要.
- 这些发现强调了需要专门的策略来处理不平衡的功能标签数据.
相关概念视频
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Conservation of Protein Domains Over Different Proteins
11.5K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.5K
Protein Families
3.4K
3.4K
Protein-protein Interfaces
13.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
13.5K
Proteomics
8.0K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
8.0K


