蛋白质语言模型的高效推断,训练和微调
Muhammed Hasan Çelik1,2, Xiaohui Xie1
1Department of Computer Science, University of California, Irvine, Irvine, CA, USA.
iScience
|October 2, 2025
概括
我们开发了一种高效的蛋白质语言模型 (ESME),可以显著加快蛋白质结构和功能预测,同时降低计算成本. 这使得对蛋白质分析的强大人工智能工具更容易被资源有限的研究人员访问.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 生命科学中的人工智能
背景情况:
- 蛋白质语言模型 (PLM) 为预测蛋白质结构和功能提供了强大的功能.
- 目前,高计算成本限制了PLM的广泛采用和可访问性.
研究的目的:
- 提高进化规模建模 (ESM) 框架的效率.
- 减少用于训练和推断大型PLM所需的计算资源.
- 为了使研究人员更容易获得先进的蛋白质分析工具.
主要方法:
- 实现了FlashAttention和序列包装,以更快地推断和减少内存使用.
- 在大参数模型中使用四位量子化来减少内存足迹.
- 使用激活检查点和DeepSpeed零卸载优化训练.
- 采用了通过适配器权重的参数效率微调.
主要成果:
- 通过FlashAttention和序列包装实现了4-9倍更快的推断和3-14倍更低的内存使用.
- 通过使用四位量子化而不会损失精度,将模型内存额外减少2-3倍.
- 通过优化技术将训练运行时间缩短6倍.
- 在蛋白质性质预测 (70%的Spearman对应点) 和功能预测 (87%的AU-PRC用于转录因子识别) 中获得了最先进的性能.
结论:
- 高效的ESM (ESME) 实施大大降低了使用PLM的计算障碍.
- 对于具有有限计算能力的学术实验室来说,ESME使访问先进的蛋白质建模工具变得民主化.
- 这项工作通过高效的AI模型促进了蛋白质科学中的更广泛的研究应用.
相关概念视频
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
Leaky Scanning
5.6K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.6K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Termination of Translation
27.4K
The large ribosomal subunit has several important structures essential to translation. These include the peptidyl transferase center (PTC) - which is the site where the peptide bond is formed - and a large, internal, water-filled tube through which the nascent polypeptide moves. This latter structure is called the Peptide Exit Tunnel, and it begins at the PTC and spans the body of the large ribosomal subunit. During translation, as the nascent polypeptide chain is synthesized, it passes through...
27.4K
Proteins: From Genes to Degradation
14.1K
Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick. Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA...
Transcription is the synthesis of RNA...
14.1K


