平整曲线 - - 如何通过小型深度突变扫描数据集获得更好的结果
Gregor Wirnsberger1, Iva Pritišanac2,3, Gustav Oberdorfer3,4
1Institute of Molecular Biosciences, University of Graz, Graz, Austria.
Proteins
|March 19, 2024
概括
这项研究引入了一种使用蛋白质结构作为神经网络输入的新方法,改进了突变效应预测. 这种方法增强了深度突变扫描 (DMS) 分析,并降低了实验成本.
科学领域:
- 生物技术是生物技术.
- 计算生物学 计算生物学
- 蛋白质工程是指蛋白质工程.
背景情况:
- 蛋白质工程通常需要通过氨基酸替代优化蛋白质特性.
- 深度突变扫描 (DMS) 是一种高通量方法,用于评估突变对蛋白质功能的影响.
- 目前用于突变预测的机器学习方法主要依赖于蛋白质序列.
研究的目的:
- 开发一种编码蛋白质结构的方法,以改善对突变影响的机器学习预测.
- 研究在较小的DMS数据集上增强神经网络训练的技术.
- 为实验DMS提供计算替代方案,用于预测突变效应.
主要方法:
- 将蛋白质结构编码为堆叠的二维接触图,捕获残留物相互作用和进化保护.
- 在使用结构输入的DMS数据集上培训基于图像分析的神经网络架构.
- 将性能与基于序列的神经网络进行比较,并探索数据增强和预训练策略.
主要成果:
- 与仅序列方法相比,编码蛋白质结构在DMS数据上显著提高了机器学习性能.
- 数据增强和预训练减少了对大型训练数据集的需求,并提高了预测准确性,特别是对于较小的数据集.
- 开发的方法实现了与在大型数据集上使用最先进的基于序列的模型可比的性能.
结论:
- 直接将结构信息纳入机器学习模型,可以更好地预测突变效应.
- 拟议的方法为实验性DMS提供了具有成本效益和效率的替代方案,加速了蛋白质工程工作流程.
- 提供了一个开源工具,使研究人员能够使用这些先进的数据分析技术.
关键词:
数据增强数据增强深度突变扫描 (deep mutational scanning) 是一种对突变进行深度扫描的方法.机器学习是机器学习.预训练的预训练蛋白质结构 蛋白质结构结构编码结构的编码.更多相关视频
09:43Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
6.3K
09:33Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
Published on: August 25, 2023
1.1K
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
