基于层次和全球特征的组合,对蛋白质序列的EC数量预测
Fan Yang1,2,3,4, Qiao-Ling Han1,2,3,4, Wen-di Zhao1,2,3,4
1School of technology, Beijing Forestry University, Beijing 100083, China.
Yi chuan = Hereditas
|August 14, 2024
概括
一种新的方法,ECPN-HFGF,通过分析蛋白质序列,准确地预测酶EC数量. 这种方法提高了对生物活动的理解,并有助于酶学研究.
科学领域:
- 生物化学 生物化学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 酶功能识别对于生物学理解和生命科学进步至关重要.
- 目前的酶EC数预测方法在准确性和蛋白质序列利用方面存在局限性.
研究的目的:
- 开发一个新的网络,ECPN-HFGF,用于准确的酶EC数量预测.
- 提高在酶功能识别中蛋白质序列信息的利用率.
主要方法:
- 利用剩余网络提取通用的蛋白质序列特征.
- 采用分层和全局特征提取模块进行全面分析.
- 集成了一个多任务学习框架,以提高预测准确度.
主要成果:
- 在EC数预测方面,ECPN-HFGF取得了卓越的表现.
- 证明了高的宏观F1 (95.5%) 和微型F1 (99.0%) 评分.
- 有效地结合了层次和全球特征,以准确预测.
结论:
- ECPN-HFGF提供了一个快速而准确的方法来预测酶EC数量.
- 该方法在预测准确性方面显著优于现有方法.
- 为酶学研究和酶工程提供了一个有效的工具.
相关概念视频
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Protein Organization
137.1K
Overview
137.1K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K


