通过参数高效的微调来民主化蛋白质语言模型
Samuel Sledzieski1,2, Meghana Kshirsagar2, Minkyung Baek3
1Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge MA 02139, USA.
bioRxiv : the preprint server for biology
|November 21, 2023
概括
像LoRA这样的参数有效微调 (PEFT) 方法被引入蛋白质组学,用于蛋白质语言模型. 这些方法显著减少了用于诸如蛋白质-蛋白质相互作用预测和同同寡合体对称性预测等任务的计算资源.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 机器学习在蛋白质学中的机器学习
背景情况:
- 大型预训练的蛋白质语言模型 (PLMs) 通过学习序列表示来彻底改变蛋白质组学.
- 精细调整特定任务的PLM是计算密集的,这对许多研究人员来说是一个障碍.
- 参数高效微调 (PEFT) 方法解决了自然语言处理中的类似挑战.
研究的目的:
- 引入和评估用于微调蛋白质组学PLM的PEFT方法.
- 评估PEFT在蛋白质-蛋白质相互作用 (PPI) 预测和同类分子对称性预测方面的表现.
- 为传统的PLM微调提供一个计算效率高的替代方案.
主要方法:
- 将LoRA (低级别调整) PEFT方法应用于PLM.
- 培训和评价PEFT模型在同类聚合物对称性预测和PPI预测任务.
- 将PEFT性能与传统的全微调和最先进的方法进行比较.
主要成果:
- 在具有显著降低记忆力和参数的同类聚合物对称性预测方面,PEFT实现了竞争性表现.
- PEFT模型在PPI预测上表现优于传统的微调,使用数量级较少的参数.
- 结PLM参数和只训练一个分类头进一步提高了PPI预测性能和参数效率.
结论:
- 在蛋白质组学中,PEFT方法提供了一种计算高效和有效的方法来适应PLM.
- 这些方法使有限的计算资源的研究人员能够获得强大的PLM调整.
- PEFT为传统微调提供了可行的替代方案,甚至在某些蛋白质经济任务中表现优于传统微调.
相关概念视频
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
56
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
56
Improving Translational Accuracy
11.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.2K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Mechanistic Models: Compartment Models in Individual and Population Analysis
43
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
43


