通过蛋白质语言模型改进了异构体蛋白质复合物的预测
Bo Chen1, Ziwei Xie2, Jiezhong Qiu3
1Department of Computer Science and Technology, Tsinghua University, Beijing, China.
Briefings in bioinformatics
|June 16, 2023
概括
使用蛋白质语言模型的新方法ESMPair识别了更好地相互作用的蛋白质同类物 (介质物) 以进行复杂结构预测. 这对AlphaFold-Multimer进行了改进,特别是在低置信度预测和真核生物复合体方面.
科学领域:
- 计算生物学 计算生物学
- 结构生物学 结构生物学
- 生物信息学是一种生物信息学.
背景情况:
- 预测蛋白质复杂结构对于理解生物功能至关重要.
- AlphaFold-Multimer是一个领先的工具,但它的准确性取决于来自交互的同类物 (互联物) 的多重序列对齐 (MSA) 的质量.
研究的目的:
- 开发一种新的方法,ESMPair,用于使用蛋白质语言模型识别高质量的互.
- 为了提高蛋白质复杂结构预测的准确性.
主要方法:
- 利用蛋白质语言模型来识别蛋白质复合体的互解物.
- 将ESMPair的互解识别与AlphaFold-Multimer的默认方法进行比较.
- 评估了MSA多样性对预测准确性的影响.
主要成果:
- 与AlphaFold-Multimer的默认方法相比,ESMPair产生了优异的互解.
- ESMPair显著提高了蛋白质复杂结构预测准确度 (+10.7% DockQ).
- 结合多种MSA生成方法进一步提高了准确性 (+22% DockQ).
- 鉴定了Interolog MSA多样性作为影响预测准确性的关键因素.
- 对于真核生物复合体,ESMPair表现出特别高的疗效.
结论:
- ESMPair提供了一种更有效的方法来识别互,从而提高了蛋白质复杂结构的预测.
- 互MSA的多样性对于准确的结构预测至关重要.
- 在ESMPair的研究中,有望推动结构生物学,特别是真核生物系统的发展.
更多相关视频
相关概念视频
Protein Complexes with Interchangeable Parts
1.9K
1.9K
Protein Complex Assembly
10.7K
Proteins can form homomeric complexes with another unit of the same protein or heteromeric complexes with different types. Most protein complexes self-assemble spontaneously via ordered pathways, while some proteins need assembly factors that guide their proper assembly. Despite the crowded intracellular environment, proteins usually interact with their correct partners and form functional complexes.
Many viruses self-assemble into a fully functional unit using the infected host cell to...
Many viruses self-assemble into a fully functional unit using the infected host cell to...
10.7K
Protein-protein Interfaces
12.6K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.6K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Conservation of Protein Domains Over Different Proteins
11.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.0K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


