用循环中的大型语言模型设计多样化和高性能蛋白质
Carlos A Gomez-Uribe1, Japheth Gado1, Meiirbek Islamov1
1Solugen, Inc., Houston, Texas, United States of America.
PLoS computational biology
|June 5, 2025
概括
本研究引入了用于蛋白质工程的机器学习,使用Seq2Fitness和为多样化和自适应序列采样 (BADASS) 进行双相回火来设计具有增强适应性的新型蛋白质序列.
科学领域:
- 蛋白质工程是指蛋白质的工程.
- 机器学习 机器学习
- 计算生物学 计算生物学
背景情况:
- 定向进化对于蛋白质工程至关重要,但通常受到采样效率的限制.
- 准确预测蛋白质健身景观并有效探索它们是关键的挑战.
研究的目的:
- 开发一种用于蛋白质工程的新型机器学习方法.
- 通过改进的健身预测和景观探索来增强序列设计.
主要方法:
- 介绍了Seq2Fitness,这是一个半监督的神经网络模型,使用蛋白质语言模型进行健身预测.
- 开发了用于多样化和自适应序列采样 (BADASS) 的双相回火,用于序列探索的优化算法.
- 集成的Seq2Fitness和BADASS用于全面的蛋白质序列设计管道.
主要成果:
- 对于未见的突变,Seq2Fitness与实验适应性 (0.55对0.34) 的相关性得到改善.
- 比起竞争的方法,BADASS产生了比较高的适应性和多样化的序列,计算成本更低.
- 100%的顶级BADASS识别的序列超过了野生类型的适应性,超过了其他方法.
结论:
- 结合Seq2Fitness和BADASS的方法显著推进了蛋白质序列设计.
- 这种方法为产生各种应用的高性能蛋白质提供了一个强大的工具.
- BADASS显示了对蛋白质以外的其他基于序列的设计问题的概括潜力.
更多相关视频
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
388
12:04Interactome-Seq: A Protocol for Domainome Library Construction, Validation and Selection by Phage Display and Next Generation Sequencing
Published on: October 3, 2018
9.1K
相关概念视频
Conservation of Protein Domains Over Different Proteins
11.5K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.5K
Conservation of Protein Domains
3.2K
3.2K
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Protein and Protein Structure
81.7K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
81.7K
