蛋白质语言模型监督 motif-scaffolding 设计使用 GPDL
Bo Zhang1, Kexin Liu1, Zhuoqi Zheng1
1State Key Laboratory of Microbial metabolism, Joint International Research Laboratory of Metabolic & Developmental Sciences, Department of Bioinformatics and Biostatistics, National Experimental Teaching Center for Life Sciences and Biotechnology, School of Life Sciences and Biotechnology, Shanghai Jiao Tong University, Shanghai, 200240, China.
International journal of biological macromolecules
|October 22, 2025
概括
语言生成蛋白质设计模型 (GPDL) 为蛋白质设计提供了一种新的方法,在生成多样化和准确的蛋白质结构方面表现优于现有的方法. 这种语言模型策略显示出创造各种应用的新型功能蛋白的前景.
科学领域:
- 计算生物学是一种计算生物学.
- 蛋白质工程是一种蛋白质工程.
- 生物信息学是一种生物信息学.
背景情况:
- 蛋白质的生物功能是由它们的3D结构决定的,特别是关键的基因残留物.
- 在动图周围生成准确和多样化的蛋白质支架仍然是现有计算方法的挑战.
- 传统的基于多个序列对齐 (MSA) 的预训方法在蛋白质结构预测和设计方面存在局限性.
研究的目的:
- 引入基于语言的生成蛋白设计 (GPDL) 模型,作为一种有效的替代方案,替代传统的基于MSA的蛋白质设计预训.
- 评估GPDL在各种基准问题中产生多样化和准确的蛋白质支架方面的表现.
- 评估GPDL的稳定性,特别是对于具有低序列相似性的孤儿蛋白质.
主要方法:
- 开发了基于语言的生成蛋白设计模型 (GPDL),以取代基于MSA的预训练.
- 为GPDL采用了一个可扩展的设计策略.
- 在24个基准蛋白质设计问题上测试了GPDL.
主要成果:
- 在24个基准问题中,GPDL成功解决了22个.
- 与射频扩散相比,GPDL产生了33.5%更多的独特可设计集群,表明了优越的多样性.
- 在各种蛋白质设计场景中证明了准确和物理可信的结构生成.
- 在孤儿蛋白质上表现出强大的稳定性,与训练数据的序列相似性有限.
结论:
- 在蛋白质设计中,GPDL有效地取代了传统的基于MSA的预训.
- 这种方法产生了准确,多样化和物理可信的蛋白质结构.
- 蛋白质语言模型对加速开发用于生物和治疗应用的新型功能蛋白具有重大前景.
相关概念视频
Assembly of Signaling Complexes
6.4K
Multiprotein signaling complexes are formed in a dynamic process involving protein-protein interactions at the cytoplasmic domain of transmembrane receptors or enzymatic and non-enzymatic proteins associated with the receptor. These complexes ensure the activation and propagation of intracellular signals that regulate cell functions.
Interaction domains in cell signaling
Interaction domains recognize exposed features of their binding partners containing post-translationally modified sequences,...
Interaction domains in cell signaling
Interaction domains recognize exposed features of their binding partners containing post-translationally modified sequences,...
6.4K
Ligand Binding and Linkage
5.4K
Allosteric proteins have more than one ligand binding site; the binding of a ligand to any of these sites influences the binding of ligands to the other sites. When a protein is allosteric, its binding sites are called coupled or linked. In the case of enzymes, the site that binds to the substrate is known as the active site and the other site is known as the regulatory site. When a ligand binds to the regulatory site, this leads to conformational changes in the protein that can influence...
5.4K
Ligand Binding and Linkage
3.9K
3.9K
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Protein-protein Interfaces
14.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.4K
GTPases and their Regulation
9.6K
Guanine nucleotide-binding proteins (G-proteins), also known as GTPases, are a superfamily of proteins that regulate many cellular processes, such as cell signaling, vesicular transport, and the regulation of cell shape and motility. Mutation or dysfunction of these proteins can lead to disease. There are around 40,000 known G-proteins that can broadly be classified into two groups ‒ small G-proteins consisting of a single domain and large multi-domain G-proteins.
Large G-proteins,...
Large G-proteins,...
9.6K


