通过trRosettaX2,AlphaFold2改善了蛋白质结构预测,并在CASP15中优化了MSA
Zhenling Peng1, Wenkai Wang2, Hong Wei2
1MOE Frontiers Science Center for Nonlinear Expectations, Research Center for Mathematics and Interdisciplinary Sciences, Shandong University, Qingdao, China.
Proteins
|August 11, 2023
概括
我们的新型管道显著提高了CASP15.15中蛋白质结构预测准确度. 通过增强多重序列对齐 (MSA) 和采用先进的建模,服务器和多元器分别获得了单元和多元器预测的顶级排名.
科学领域:
- 计算生物学 计算生物学
- 结构生物学 结构生物学
- 生物信息学是一种生物信息学.
背景情况:
- 准确的蛋白质结构预测对于理解生物功能和疾病机制至关重要.
- 蛋白质结构预测的批判性评估 (Critical Assessment of Protein Structure Prediction,CASP) 是一个社区范围的实验,用于评估蛋白质结构预测方法的准确性.
- 深度学习的进步已经彻底改变了蛋白质结构的预测,但对于复杂的目标,如多元体,仍然存在挑战.
研究的目的:
- 在CASP15竞赛中展示我们新型蛋白质结构预测管道的性能.
- 为了评估我们的方法在预测单体和多体蛋白质结构方面的有效性.
- 确定有助于预测准确性的关键因素和需要进一步发展的领域.
主要方法:
- 开发了一个复杂的管道,利用互补的序列数据库和先进的搜索算法来生成高质量的多重序列对齐 (MSA).
- 使用trRosettaX2和AlphaFold2进行单体结构预测 (Yang-Server) 和AlphaFold-Multimer进行多体结构预测 (Yang-Multimer).
- 将预测结果与默认的AlphaFold2和AlphaFold-Multimer实现进行了比较.
主要成果:
- 服务器在单体结构预测方面获得了最高排名,平均TM分数为0.876,超过了默认AlphaFold2 (0.798).
- -Multimer在多重结构预测方面排名第四,平均DockQ得分为0.464,超过默认的AlphaFold-Multimer (0.389).
- 改进归因于增强的MSA,针对大型目标的代建模,以及单体和多体预测策略之间的相互作用.
结论:
- 我们的管道证明了单体和多体蛋白质结构预测准确性的显著改进.
- 增强的MSA和先进的建模技术是预测性能的关键驱动因素.
- 孤儿蛋白和复杂多元体的结构预测仍然是一个具有挑战性的领域,需要未来的突破.
相关概念视频
Protein Organization
6.6K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.6K
Protein Folding
118.5K
Overview
118.5K
Protein Folding Quality Check in the RER
3.7K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.7K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Protein and Protein Structure
79.7K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
79.7K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


