大型蛋白质语言模型的参数高效微调可以改善信号预测
Shuai Zeng1, Duolin Wang1, Lei Jiang1
1Department of Electrical Engineering and Computer Science, Christopher S. Bond Life Sciences Center, University of Missouri, Columbia, Missouri 65211, USA.
Genome research
|July 26, 2024
概括
我们开发了PEFT-SP,这是一个使用参数有效微调 (PEFT) 改进信号 (SP) 预测的新框架. 这种方法通过利用蛋白质语言模型显著提高准确性,特别是对于有限的数据.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 蛋白质的结构和功能.
背景情况:
- 信号 (SP) 对于细胞内的蛋白质定位至关重要.
- 大型蛋白质语言模型 (PLM) 为SP预测提供了新的途径,特别是对于数据稀缺的类别.
- 有效利用PLM需要先进的微调策略.
研究的目的:
- 引入PEFT-SP,一个参数高效的微调框架,用于增强的信号预测.
- 利用PLM中编码的进化信息来改进SP识别.
- 在SP预测的背景下评估不同PEFT技术的性能.
主要方法:
- 低级适应 (LoRA) 与ESM-2蛋白语言模型的整合.
- 在ESM-2框架内,应用快速调整和适配器调整作为替代PEFT方法.
- 与使用马修斯相关系数 (MCC) 的最先进方法对比PEFT-SP性能的比较分析.
主要成果:
- 使用LoRA的PEFT-SP实现了显著的改进,在小型培训样本的SP中获得了87.3%的MCC增长,整体MCC增长6.1%.
- 采用适配器调的PEFT-SP表现出显著的收益,在小型培训样本的SP中,MCC高达28.1%,整体为3.8%.
- 与适配器调相比,LoRA表现出更高的计算效率和更低的内存需求.
结论:
- PEFT-SP有效地提高了信号预测的准确性,特别是对于代表性不足的类别.
- 洛拉 (LoRA) 提出了一种计算效率高且有效的PEFT方法,用于调整大型PLM用于SP预测.
- PEFT-SP框架促进了强大的蛋白质模型的适应,用于关键的生物任务,如信号的识别.
相关概念视频
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Improving Translational Accuracy
9.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.7K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K


