EvoRator2:使用深度学习基于蛋白质结构信息预测特定站点的氨基酸替代.
Natan Nagar1, Jérôme Tubiana2, Gil Loewenthal1
1The Shmunis School of Biomedicine and Cancer Research, George S. Wise Faculty of Life Sciences, Tel Aviv University, Tel Aviv 69978, Israel.
Journal of molecular biology
|June 25, 2023
概括
EvoRator2使用蛋白质结构预测耐受性氨基酸,有助于对蛋白质的研究,其序列数据有限. 这种深度学习工具可以改进突变效应的预测,特别是对于孤儿和de novo蛋白质.
科学领域:
- 计算生物学是一种计算生物学.
- 结构生物学是结构生物学.
- 分子进化是分子进化的过程.
背景情况:
- 多个序列对齐 (MSAs) 推断耐受性氨基酸,但对于少数同类蛋白质而言是有限的.
- 孤儿和de novo设计的蛋白质缺乏足够的同源序列用于传统的基于MSA的分析.
研究的目的:
- 开发一种深度学习算法 (EvoRator2),只使用蛋白质结构信息来预测耐受性氨基酸.
- 为了应对具有有限或没有同源序列的蛋白质分析的挑战.
主要方法:
- 埃沃拉托2在超过15000个蛋白质结构上受过训练.
- 该算法根据来自原子坐标文件的结构信息预测耐受性氨基酸.
- 使用位置加权评分矩阵 (PSSM) 和深度突变扫描 (DMS) 实验来评估性能.
主要成果:
- 在预测PSSM方面,EvoRator2显示了令人满意的结果.
- 它在DMS实验中预测突变效应方面取得了近乎最先进的性能,在某些目标上表现优于现有的方法.
- 将EvoRator2与基于MSA的方法结合起来,提高了DMS实验的预测准确性和稳定性.
结论:
- 通过单独使用蛋白质结构,EvoRator2有效地预测耐受性氨基酸替代.
- 该工具对于研究孤儿蛋白和新设计的蛋白质非常有价值.
- 埃沃拉托网络服务器可用于更广泛的应用.
相关概念视频
Conservation of Protein Domains Over Different Proteins
11.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.0K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K
Conservation of Protein Domains
3.1K
3.1K
Amino acids
89.3K
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible...
89.3K
Predicting Reaction Outcomes
8.5K
Kinetics describes the rate and path by which a reaction occurs. In contrast, thermodynamics deals with state functions and describes the properties, behavior, and components of a system. It is not concerned with the path taken by the process and cannot address the rate at which a reaction occurs. Although it does provide information about what can happen during a reaction process, it does not describe the detailed steps of what appears on an atomic or a molecular level. On the other hand,...
8.5K


