大规模预测蛋白质功能通过异质特征的融合.
Rongtao Zheng1, Zhijian Huang1, Lei Deng1
1School of Computer Science and Engineering, Central South University, 410000 Changsha, China.
Briefings in bioinformatics
|July 4, 2023
概括
PredGO通过将AlphaFold结构预测与非结构数据集成来增强蛋白质功能注释. 这种方法显著提高了基因本体学 (GO) 对蛋白质的功能预测的准确性和覆盖范围.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 结构生物学 结构生物学
背景情况:
- 蛋白质序列和结构数据的快速增长需要用于函数注释的自动化方法.
- 现有的用于蛋白质功能预测的计算方法在准确性和覆盖范围上有局限性.
- 蛋白质功能的实验性确定无法扩展到大量可用的数据.
研究的目的:
- 开发一种大规模的计算方法,用于对蛋白质的基因本体学 (GO) 函数进行注释.
- 为了利用AlphaFold预测的蛋白质结构以及非结构性线索来改善功能预测.
- 创建一个强大的和准确的方法,用于自动化蛋白质功能注释.
主要方法:
- 使用AlphaFold预测了三维结构信息作为一个关键特征.
- 集成的非结构性线索,如序列同质性,蛋白质与蛋白质相互作用和基因共同表达.
- 采用预先训练的语言模型,几何向量感知子,以及用于特征提取和融合的注意力机制.
主要成果:
- 与最先进的方法相比,PredGO方法在预测GO函数方面表现优越.
- 在蛋白质功能预测的覆盖范围和准确性方面取得了显著的改进.
- 成功注释了超过205,000个人类UniProt条目,大约90%基于预测的结构.
结论:
- PredGO有效地集成结构和非结构数据,以准确和全面的蛋白质功能注释.
- 该方法解决了现有方法的局限性,为大规模注释提供了可扩展的解决方案.
- 一个Web服务器和数据库是公开的,促进PredGO工具的更广泛的应用.
相关概念视频
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Tagging and Fusion Proteins
6.8K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.8K
Protein-protein Interfaces
12.6K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.6K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Protein Families
15.5K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.5K


