利用单细胞大型语言模型的力量,使用scPEFT进行参数高效微调
Fei He1, Ruixin Fei1, Jordan E Krull2,3
1Department of Electrical Engineering and Computer Science, Bond Life Sciences Center, University of Missouri, Columbia, MO, 65211, USA.
Research square
|May 2, 2025
概括
我们开发了单细胞参数效率微调 (scPEFT) 来改进单细胞大语言模型 (scLLM). scPEFT提高了模型的适应性和可访问性,用于研究人员使用有限的数据.
科学领域:
- 计算生物学 计算生物学
- 基因组学就是基因组学.
- 人工智能的人工智能
背景情况:
- 单细胞大语言模型 (scLLM) 从单细胞地图集提供了强大的洞察力.
- 然而,scLLMs在语境之外的应用中表现出局限性,导致不可靠的零射击预测.
研究的目的:
- 引入一个新的框架,单细胞参数效率微调 (scPEFT),以提高scLLM的性能.
- 提高scLLM的适应性和可访问性,用于各种生物研究.
主要方法:
- 通过将低维适配器集成到 scLLM 中来实现 scPEFT.
- 结骨干模型,仅更新适配器参数,以有效地适应有限数据的任务.
- 参数调整减少了96%以上,GPU内存使用量减少了50%以上.
主要成果:
- 与零射击和传统微调方法相比,scPEFT在各种数据集中表现出卓越的性能.
- 成功应用于疾病特异性,跨物种和低特征细胞群体分析.
- 注意力机制分析确定了与COVID相关的基因和新的血细胞亚群.
结论:
- scPEFT提供了一种高效且易于使用的解决方案,用于使scLLM适应特定的生物任务.
- 该框架增强了scLLM用于一般单细胞分析的实用性,特别是在资源有限的环境中.
- scPEFT促进了特定条件的生物学解释和发现.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


