使用深度生成模型寻找具有所需性质的蛋白质变体.
Yan Li1, Yinying Yao2,3, Yu Xia1
1School of Information, Yunnan Normal University, Kunming, China.
BMC bioinformatics
|July 21, 2023
概括
本研究引入了一个时间变异自编码器 (T-VAE) 模型,通过改善较长的氨基酸序列的表示和产生更相似的蛋白质变异来增强蛋白质工程. 与基线模型相比,T-VAE模型在预测蛋白质适应性和序列身份方面表现出卓越的性能.
科学领域:
- 生物技术是生物技术.
- 计算生物学 计算生物学
- 蛋白质工程是指蛋白质工程.
背景情况:
- 蛋白质工程旨在提高各种应用的蛋白质功能.
- 深度学习模型捕获蛋白质序列特征,但与远程依赖性作斗争.
- 现有的生成模型需要改进,以更长的序列中代表氨基酸位点之间的关系.
研究的目的:
- 开发一种深度学习模型,以改善对较长蛋白序列的表示学习.
- 为了增强生成的蛋白质序列和原始序列之间的相似性.
- 为了利用潜在空间中的蛋白质序列的位置关系来发现变体.
主要方法:
- 提出了一个时间变化自编码器 (T-VAE) 模型,包括一个编码器和一个解码器.
- 在编码器中利用扩展因果卷积来扩展受体场,以获得更好的长序编码.
- 解码器生成了与原始蛋白质序列非常相似的变体.
主要成果:
- 与其他模型相比,T-VAE在预测蛋白质适应性方面实现了更高的人相关系数和更低的平均绝对偏差.
- 对于较长的蛋白质序列,证明了优越的表示学习.
- 与基线模型相比,生成和输入数据之间的序列一致性得到了12.9%的改善.
结论:
- 对于较长的蛋白质序列,T-VAE在表示学习方面表现出增强的能力.
- 该模型显示,在产生具有高序列同一性的蛋白质变体方面具有显著的优势.
- T-VAE提供了一种有前途的方法,用于发现具有改善功能性质的新型蛋白质变体.
相关概念视频
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Complexes with Interchangeable Parts
2.6K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.6K
Conservation of Protein Domains
3.1K
3.1K
Protein-protein Interfaces
12.6K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.6K


