蛋白质化温度的监督学习:跨物种与特定物种的预测
Sebastián García López1, Jesper Salomon2, Wouter Boomsma1
1Department of Computer Science-DIKU, University of Copenhagen, Copenhagen, Denmark.
Proteins
|July 14, 2025
概括
预测蛋白质融化温度对于蛋白质工程至关重要. 跨物种训练模型高估了性能,无法准确预测工程变体的变化或识别热稳定蛋白质.
科学领域:
- 生物化学 生物化学
- 结构生物学 结构生物学
- 计算生物学 计算生物学
背景情况:
- 蛋白质化温度 (Tm) 是蛋白质稳定性的关键指标,对于蛋白质工程应用,如酶发现和优化至关重要.
- 自然蛋白质的Tm值的大数据集使得预测模型的开发成为可能,报告的相关系数高 (Spearman rho).
- 这些高相关性表明,可以准确地预测工程变异中的Tm变化,并识别热稳定蛋白质,尽管实际结果往往不够理想.
研究的目的:
- 调查报告的高相关性得分与蛋白质化温度预测模型的实际性能之间的差异.
- 评估跨物种培训与预测蛋白质融化温度的特定物种模型之间的有效性.
- 评估转移学习方法对不同物种Tm预测的有用性.
主要方法:
- 对跨物种数据的Spearman rho相关性进行分析,以评估其对预测性能的代表性.
- 跨物种培训与特定物种模型培训的比较,使用四种转移学习方法和微调程序.
- 评估模型在预测工程变体化温度变化和识别自然恒温蛋白质方面的性能.
主要成果:
- 斯皮尔曼罗对跨物种数据提供了一个过于乐观的预测性能,主要反映全球氨基酸组成差异,而不是特定的遗传变异效应.
- 跨物种培训在培训特定物种模型预测融化温度方面没有表现出一致的好处.
- 目前用于预测蛋白质化温度的监督模型的表现明显低于以前文学指标所建议的表现.
结论:
- 通常使用的斯皮尔曼rho度量可以误导,当应用到跨物种数据蛋白质化温度预测.
- 跨物种的学习转移用于预测蛋白质化温度仍然是一个重大挑战.
- 需要进行进一步的研究,以开发更强大,更准确的模型来预测蛋白质稳定性,特别是在工程变体和各种物种中.
相关概念视频
Conservation of Protein Domains Over Different Proteins
11.4K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.4K
Conserved Binding Sites
4.4K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.4K
Multi-species Conserved Sequences
4.3K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.3K


