MSNGO:基于3D蛋白质结构和网络传播的多种蛋白质功能注释
Beibei Wang1, Boyue Cui1, Shiqu Chen1
1School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, Guang Dong 518055, China.
Bioinformatics (Oxford, England)
|May 6, 2025
概括
通过整合蛋白质结构和网络传播,MSNGO模型增强了多种蛋白质功能预测. 这种方法改善了对蛋白质注释有限的物种的跨物种标签传播.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 结构生物学 结构生物学
背景情况:
- 使用AlphaFold2预测结构,蛋白质功能预测的准确性已经提升,超过仅序列的方法.
- 多种蛋白质功能预测在数据集成和不同物种之间的知识转移方面面临挑战.
- 对稀有标注物种进行有效的跨物种标签传播仍然是一个关键的挑战.
研究的目的:
- 提出一种新型模型,MSNGO (多种蛋白质结构和预测GO术语的网络),用于改进多种蛋白质功能预测.
- 整合结构特征和网络传播,以解决当前多种方法的局限性.
- 为了增强具有稀疏蛋白质注释的物种的跨物种标签传播.
主要方法:
- 使用图形表示学习从蛋白质结构联系图中提取氨基酸特征.
- 使用图形卷积聚合模块来导出蛋白质水平的结构特征.
- 整合ESM-2的序列特征,并在异质网络中应用网络传播.
主要成果:
- 该MSNGO模型整合了结构和序列特征,用于增强蛋白质功能预测.
- 网络传播有效地汇总信息并更新异质网络中的节点表示.
- 与现有的多种方法相比,MSNGO表现出优异的性能,仅依赖于序列特征和蛋白质-蛋白质网络.
结论:
- 整合结构特征显著提高了多种蛋白质功能预测的准确性.
- MSNGO为跨物种标签的传播提供了一个有效的解决方案,特别是对于标注有限的物种.
- 开发的模型通过利用结构和网络信息来推进多种蛋白质功能的预测领域.
相关概念视频
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Protein-protein Interfaces
12.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.4K
Protein Organization
6.0K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.0K
Conserved Binding Sites
4.1K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.1K
Genome Annotation and Assembly
18.7K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.7K
Protein Families
15.1K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.1K


