收获LLM修剪的果实:朝着小型语言模型,以实现高效的非编码变体效应预测
Megha Hegde1, Jean-Christophe Nebel2, Farzana Rahman1
1School of Computer Science and Mathematics, Kingston University, London KT1 2EE, UK.
层修剪使大型基因组语言模型更有效地进行变异预测. 删除冗余层可以降低计算需求,而不会牺牲准确性,改善非编码变体解释.
科学领域:
- 基因组学就是基因组学.
- 计算生物学 计算生物学
- 人工智能的人工智能
背景情况:
- 解释基因变异对于精准医学至关重要.
- 大型基因组语言模型 (LLM) 由于计算缩放而难以预测非编码变体.
- 在自然语言处理中成功的层修剪可以优化LLMs.
研究的目的:
- 系统地评估基因组LLM (DNABERT 2,核酸转换器) 中每个转换器层对变体预测的贡献.
- 通过去除非关键层来开发更简单,更高效的计算效率的LLM.
- 在非编码变体效应预测基准上评估修剪后的LLM的性能.
主要方法:
- 在DNABERT 2和核酸转换器模型中进行系统的层切除.
- 基于性能变化构建层层重要性概况.
- 在Enformer eQTL因果变异数据集上的微调修剪和完整模型.
- 比较性能指标 (准确度,AUC) 和资源使用 (训练时间,内存).
主要成果:
- 层的重要性在不同模型之间有很大差异,有些层可以在最小的性能损失下移除.
- 经过微调后,修剪后的模型实现了与完整模型相比的精度和AUC.
- 修剪后的模型显示,训练时间和记忆需求大幅减少.
结论:
- 层 wise 修剪是一种有效的策略,用于创建紧和高效的基因组LLMs.
- 修剪后的LLM保持预测能力,同时显著降低计算需求.
- 这种方法提高了用于研究和临床应用的大规模非编码变异分析的可访问性.
更多相关视频
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
相关概念视频
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Improving Translational Accuracy
Improving Translational Accuracy
lncRNA - Long Non-coding RNAs
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:
