蛋白质适应性预测受到语言模型,合奏学习和采样方法的相互作用的影响
Mehrsa Mardikoraem1,2, Daniel Woldring1,2
1Department of Chemical Engineering and Materials Science, Michigan State University, East Lansing, MI 48824, USA.
Pharmaceutics
|May 27, 2023
概括
机器学习 (ML) 通过分析序列数据来增强蛋白质设计. 这项研究通过比较采样技术和序列表示来优化蛋白质工程的ML模型,改善结合亲和力和稳定性预测.
科学领域:
- 蛋白质工程是指蛋白质工程.
- 计算生物学 计算生物学
- 机器学习应用 机器学习应用
背景情况:
- 高通量测序和机器学习 (ML) 推进了新型蛋白质设计.
- 蛋白质设计的ML中的挑战包括不平衡的数据集和序列表示选择.
- 需要指导训练和评估ML方法对数据的测序.
研究的目的:
- 为蛋白质工程的测试标记数据集应用ML提供一个框架.
- 评估采样技术和蛋白质序列表示对预测任务的影响.
- 改进使用ML的结合亲和度和热稳定性预测.
主要方法:
- 对比了蛋白质序列表示的One-Hot,物理化学,UniRep (下一个令牌预测) 和ESM (掩盖令牌预测).
- 评估采样技术,包括合成少数超样本技术 (SMOTE) 和低样本技术.
- 使用多重标准决策分析 (MCDA) 与TOPSIS和权衡进行方法排名.
主要成果:
- 在使用One-Hot,UniRep和ESM序列表示时,SMOTE的表现优于低采样.
- 组合学习在亲和数据集上的预测性能提高了4% (F1分数=97%).
- 在稳定性预测方面,ESM代表实现了高准确度 (F1得分=92%).
结论:
- 该研究为蛋白质工程中的ML应用提供了一个强大的框架.
- 优化的采样和序列表示方法提高了蛋白质设计的ML模型性能.
- 这项工作有助于更准确地预测蛋白质结合亲和力和热稳定性.
关键词:
在MCDA中,MCDA是MCDA.这里是TOPSIS的地图.嵌入式 嵌入式 嵌入式组合学习组合学习不平衡的测试标记数据集.机器学习是机器学习.蛋白质健身预测 蛋白质健身预测采样方法 采样方法序列表示表示 序列表示更多相关视频
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
68
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
68
Inclusive Fitness
36.1K
Most altruistic behavior—in which one animal helps another at a cost to themselves—occurs between relatives. Scientists think these altruistic behaviors evolved because they increase the inclusive fitness of the animal providing help.
36.1K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
88
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
88
Model Approaches for Pharmacokinetic Data: Physiological Models
80
Physiological models in pharmacokinetics are instrumental in understanding the distribution and elimination of drugs within the body. These models describe the drug concentration within target organs, influenced by factors such as drug uptake, tissue volume, and blood flow. Drug uptake is governed by the partition coefficient, which signifies the drug concentration ratio in tissue to that in the blood. The blood flow rate to a specific tissue is expressed as Qt, and the rate of change in tissue...
80
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K


