蛋白质语言模型和基于结构的机器学习用于预测蛋白质激酶中的体结合位:基于能源景观编码的挫折感的可解释的人工智能框架
bioRxiv : the preprint server for biology
|January 16, 2026
概括
使用人工智能预测蛋白质激酶中的全结合位是具有挑战性的,因为它们的隐秘性质. 蛋白质丧分析显示,异位与正位不同,具有中性突变约束,解释了AI性能差异.
科学领域:
- 生物化学和结构生物学
- 计算生物学和化学信息学
- 药物发现和药物化学
背景情况:
- 识别全结合点对于药物发现至关重要,但仍然是一个重大挑战,特别是在蛋白质激酶中.
- 阿洛斯特遗址通常结构上是神秘的,进化上是未保存的,人口稀少,因此难以预测.
- 现有的计算方法难以准确预测这些具有挑战性的地点.
研究的目的:
- 系统地分析精细调整的蛋白质语言模型 (PLM) 和基于结构的方法 (P2Rank) 的性能,用于预测人体激酶中的正和结位.
- 通过局部丧分析,研究 ортостерик和全位预测之间的性能差异的机制基础.
- 在结合部位预测中重新定义AI性能,以反映功能设计和蛋白质丧.
主要方法:
- 使用了453个人类激酶-连接体综合体的精选数据集,涵盖五个抑制剂类别.
- 一个预训练的蛋白质语言模型 (ESM2-650M) 进行了微调,以进行结合部位预测.
- 基于序列的PLM和基于结构的P2Rank被用于网站识别.
- 集成了大规模的局部丧分析来解释预测差异.
主要成果:
- 无论是PLM还是P2Rank,都在正经位点实现了高性能 (AUPR = 0.64-0.76).
- 对于全位的PLM性能显著下降 (AUPR = 0.06),尽管排名能力中等 (AUROC = 0.70).
- 局部丧分析显示,orthosteric位点富含最小丧残留物,而allosteric位点表现出中性突变丧,表明进化的宽容性.
结论:
- 人工智能模型在预测全与正激酶位点方面的性能差距与它们独特的潜在生物物理性质有关,特别是突变约束.
- 蛋白质丧作为一个可解释的AI框架,以合理化AI性能在结合部位预测.
- 了解这些差异是改善基于结构的药物发现的关键.
相关概念视频
Protein-protein Interfaces
14.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.4K
Protein-Protein Interfaces
4.4K
4.4K
Ligand Binding Sites
14.9K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
14.9K
Ligand Binding Sites
8.6K
8.6K
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Allosteric Proteins-ATCase
6.5K
Binding sites linkages can regulate a protein's function. For example, enzyme activity is often regulated through a feedback mechanism where the end product of the biochemical process serves as an inhibitor.
Aspartate transcarbamoylase (ATCase) is a cytosolic enzyme that catalyzes the condensation of L-aspartate and carbamoyl phosphate to N-carbamoyl-L-aspartate. This reaction is the first step in pyrimidine biosynthesis. UTP and CTP, the end products of the pyrimidine synthesis...
Aspartate transcarbamoylase (ATCase) is a cytosolic enzyme that catalyzes the condensation of L-aspartate and carbamoyl phosphate to N-carbamoyl-L-aspartate. This reaction is the first step in pyrimidine biosynthesis. UTP and CTP, the end products of the pyrimidine synthesis...
6.5K


