Rprot-Vec:一种用于快速计算蛋白质结构相似性的深度学习方法
1Department of Graduate School of Frontier Sciences, The University of Tokyo, Tokyo, Japan.
BMC bioinformatics
|July 10, 2025
概括
Rprot-Vec是一种新的深度学习模型,可以预测蛋白质的结构相似性,并仅使用氨基酸序列检测同质性. 这种快速轻便的工具有助于理解蛋白质的功能,并使新的生物学发现.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 结构生物信息学 结构生物信息学
背景情况:
- 预测蛋白质结构相似性和检测同源序列对于推断蛋白质功能至关重要.
- 当前的方法往往依赖于3D结构数据,而这不是普遍可用的.
- 需要有效的基于序列的方法来进行大规模的结构相似性预测.
研究的目的:
- 开发一种深度学习模型,仅使用初级序列数据来预测蛋白质结构相似性和同质检测.
- 创建一个快速轻量级的解决方案,克服传统的3D结构依赖方法的局限性.
主要方法:
- 该Rprot-Vec (快速蛋白质向量) 模型集成了双向GRU和多尺度CNN层.
- 基于ProtT5的编码用于高效的序列表示.
- 该模型使用精心策划的数据集进行了训练和评估.
主要成果:
- 对于同类蛋白质,Rprot-Vec实现了65.3%的准确相似性预测率 (TM-score > 0.8).
- 该模型在所有TM得分间隔中显示了0.0561的平均预测误差.
- 在所有测试的场景中,Rprot-Vec在所有测试场景中都超过了现有的TM-vec基线,尽管参数较少.
结论:
- Rprot-Vec为结构相似性预测提供了一种快速有效的基于序列的方法.
- 该模型在蛋白质同质检测,结构功能推断和药物重定位方面具有广泛的应用.
- 开源可用性和发布的数据集将促进社区采用和进一步研究.
相关概念视频
Protein Organization
7.3K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
7.3K
Protein and Protein Structure
81.5K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
81.5K
Conservation of Protein Domains Over Different Proteins
11.4K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.4K
Protein-protein Interfaces
13.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
13.4K
Protein and Protein Structures
10.9K
10.9K
Conserved Binding Sites
4.4K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.4K


