提高蛋白质语言模型的效率,使用最小的湿实验室数据,通过几次射击学习来提高效率
Ziyi Zhou1,2, Liang Zhang1, Yuanxi Yu1
1School of Physics and Astronomy, Shanghai Jiao Tong University, Shanghai, 200240, China.
Nature communications
|July 2, 2024
概括
FSFP是一种新的训练策略,它增强了蛋白质语言模型的适应性预测,即使数据有限. 这种方法通过提高模型准确性和可解释性来帮助人工智能引导的蛋白质工程.
科学领域:
- 蛋白质工程是一种蛋白质工程.
- 计算生物学是一种计算生物学.
- 机器学习 机器学习
背景情况:
- 精确的蛋白质健身景观建模对于蛋白质工程至关重要.
- 预训练的蛋白质语言模型 (PLMs) 在健康预测方面表现出色,但在准确性和可解释性方面存在局限性.
- 传统的监督模型需要大量的数据,但往往无法获得这些数据.
研究的目的:
- 引入FSFP,一种有效的训练策略,用于在蛋白质适应性预测数据稀缺的情况下优化PLM.
- 提高PLM在预测蛋白质适应性的准确性和可解释性.
主要方法:
- FSFP 结合了元转移学习,学习排名和参数高效微调.
- 该策略优化了PLM,使用最小的标记单站位突变数据.
- 使用87个深度突变扫描数据集进行评估,用于in silico基准.
主要成果:
- FSFP显著提高了各种PLM的性能,其标记数据有限.
- 在in silico基准中表现出优于无监督和监督基线方法的优越性.
- 成功地将FSFP应用于工程 Phi29 DNA聚合酶,在湿实验室实验中将阳性率提高了25%.
结论:
- FSFP有效地解决了在训练蛋白质语言模型进行健身预测时的数据稀缺性挑战.
- 这种方法显示了推进人工智能引导蛋白质工程的巨大潜力.
- FSFP提供了一种改善蛋白质设计和工程工作流程的实用解决方案.
相关概念视频
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


