面具反向折叠与序列转移用于蛋白质表示学习
Kevin K Yang1, Niccolò Zanichelli2, Hugh Yeh3
1Microsoft Research, 1 Memorial Drive, Cambridge, MA, USA.
Protein engineering, design & selection : PEDS
|October 26, 2023
概括
这项研究引入了一种新的蛋白质语言模型,该模型结合了序列和结构信息,以改进蛋白质工程. 该模型通过利用现有的序列和结构数据来增强蛋白质功能预测.
科学领域:
- 计算生物学是一种计算生物学.
- 蛋白质工程是一种蛋白质工程.
- 机器学习是机器学习.
背景情况:
- 蛋白质序列的自我监督学习在功能和健康预测方面取得了最先进的结果.
- 仅序列模型忽略了结构信息,而反向折叠方法不能充分利用没有已知的结构的可用序列.
研究的目的:
- 开发和评估一个集成结构信息的蒙面逆折叠蛋白语言模型.
- 调查组合基于序列和基于结构的预训练对蛋白质工程任务的影响.
主要方法:
- 一个蒙面的反向折叠蛋白语言模型被训练成结构图神经网络.
- 该模型在预训练期间重建受损的蛋白质序列,这些蛋白质序列在预训练期间受到了脊柱结构的条件.
- 来自仅序列蛋白语言模型的输出被用作反向折叠模型的输入,以评估性能改进.
主要成果:
- 拟议的模型在结合仅序列模型的模型输出时,证明了改进的预训练困难性.
- 对下游蛋白质工程任务的评估表明,使用来自实验或预测结构的结构信息的好处.
- 该研究分析了归因于整合结构数据的绩效增长.
结论:
- 将结构信息集成到蛋白质语言模型中可以显著提高蛋白质工程任务的性能.
- 开发的蒙面反向折叠模型提供了一种新的方法,可以利用序列和结构数据.
- 这项工作推进了机器学习在理解和工程蛋白质功能的应用.
相关概念视频
Protein Folding
118.3K
Overview
118.3K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Conservation of Protein Domains
3.1K
3.1K
Protein Folding Quality Check in the RER
3.7K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.7K
Molecular Chaperones and Protein Folding
18.0K
The native conformation of a protein is formed by interactions between the side chains of its constituent amino acids. When the amino acids cannot form these interactions, the protein cannot fold by itself and needs chaperones. Notably, chaperones do not relay any additional information required for the folding of polypeptides; the native conformation of a protein is determined solely by its amino acid sequence. Chaperones catalyze protein folding without being a part of the folded protein.
The...
The...
18.0K
Protein and Protein Structures
10.5K
10.5K


