VEFill:一种模型,用于在蛋白质域中准确和可概括的深度突变扫描得分计算
Polina V Polunina1, Wolfgang Maier1, Alan F Rubin2,3
1Bioinformatics Group, Department of Computer Science, University of Freiburg, Freiburg, Germany.
bioRxiv : the preprint server for biology
|June 4, 2025
概括
使用渐变增强模型,VEFill有效地归因了缺失的深度突变扫描 (DMS) 评分,从而改善了变体解释. 这种方法优于基于稳定性的测试和有限的数据,有助于蛋白质工程和变异效应研究.
科学领域:
- 计算生物学是一种计算生物学.
- 蛋白质工程是一种蛋白质工程.
- 生物信息学是一种生物信息学.
背景情况:
- 深度突变扫描 (DMS) 试验系统地评估氨基酸替代对蛋白质功能的影响.
- 技术限制往往导致DMS数据集中的变体覆盖不完整,阻碍了变体解释.
- 需要计算方法来解决DMS实验中缺少的数据.
研究的目的:
- 开发和验证VEFill,一种机器学习模型,用于归纳缺失的DMS分数.
- 整合多样化的生物特征,以准确的归算.
- 评估VEFill在不同试验类型和数据稀疏性级别的性能.
主要方法:
- 开发了VEFill,这是一个在Human Domainome 1数据集上训练的梯度提升模型.
- 集成功能:ESM-1v序列嵌入,进化保存 (EVE分数),替代矩阵和物理化学描述器.
- 使用R平方和皮尔森相关性评估性能,对未见的蛋白质和不同试验类型进行测试概括.
主要成果:
- 对于DMS得分的归算,VEFill取得了强大的预测性能 (R2 = 0.64,Pearson r = 0.80).
- 在基于稳定性的数据集中对未见的蛋白质进行可靠的概括,在基于活动的测试中表现较差.
- 一个计算效率高的两个特征模型 (ESM-1v嵌入和平均DMS得分) 的性能与完整模型相比.
结论:
- VEFill提供了一个可解释和可扩展的框架,用于赋值DMS得分,特别有效用于以稳定性为重点和稀疏数据场景.
- 该模型有助于系统的突变优先级,并可以帮助设计高效的实验图书馆,用于变异效应研究.
- 虽然有效,但对复杂蛋白质的真实零射击预测仍然是一个挑战.
关键词:
在DMS中,分数的归算是DMS的分数.深度突变扫描 (deep mutational scanning) 是一种对突变进行深度扫描的方法.功能集成 功能集成 功能集成机器学习是机器学习.蛋白质的稳定性 蛋白质的稳定性序列嵌入式的嵌入式变量效应预测器变量效应预测器零射击学习的学习更多相关视频
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
1.4K
06:41In Vivo Functional Study of Disease-associated Rare Human Variants Using Drosophila
Published on: August 20, 2019
14.2K
相关概念视频
Conservation of Protein Domains Over Different Proteins
14.1K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.1K
Conservation of Protein Domains
4.0K
4.0K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K
