弥合差异移动,Log S和Log P之间的差距,使用机器学习和SHAP分析
Cailum M K Stienstra1, Christian Ieritano1, Alexander Haack1
1Department of Chemistry, University of Waterloo, Waterloo, Ontario N2L 3G1, Canada.
Analytical chemistry
|June 29, 2023
概括
用差分移动谱法 (DMS) 数据训练的机器学习模型准确地预测药物特性,如水溶性 (log S) 和疏水性 (log P). 这种方法为用于候选药物查的纯结构模型提供了有价值的替代方案.
科学领域:
- 计算化学和化学信息学
- 分析化学 分析化学
- 在化学科学中的机器学习
背景情况:
- 水溶性 (log S) 和水-醇分区系数 (log P) 是药物开发和环境命运评估的关键物理化学性质.
- 准确预测这些特性对于有效选候选药物和理解化学物质运输至关重要.
- 现有的方法通常依赖于大型数据集或纯粹基于结构的模型,这些模型可能有局限性.
研究的目的:
- 开发和评估机器学习 (ML) 框架,用于使用差分流动性谱度 (DMS) 数据预测水溶性 (log S) 和疏水性 (log P).
- 为了评估ML模型的可解释性,使用SHapley添加式扩展 (SHAP) 分析.
- 为了比较基于DMS的模型的性能,有或没有额外的结构描述符.
主要方法:
- 在微溶解环境中进行了差分流动性谱法 (DMS) 实验,以生成离子流动性/DMS数据 (例如碰撞截面,分散曲线).
- 机器学习回归器和集合堆叠被用来构建对log S和log P的预测模型.
- 为了获得333个分析品的参考日志S和日志P值,使用了OPERA包,用于模型解释性,使用了SHAP分析.
主要成果:
- 基于DMS的回归模型实现了log S和log P预测的R2 = 0.67,其中RMSE值分别为1.03 ± 0.10和1.20 ± 0.10.
- SHAP分析表明,气相聚类是log P预测的一个重要因素.
- 整合结构描述符提高了模型性能,对log S产生RMSE=0.84和R2=0.78,对log P产生RMSE=0.83和R2=0.84.
结论:
- 差分移动谱法 (DMS) 数据,当与机器学习一起使用时,为预测关键物理化学性质 (如log S和log P) 提供了一种有价值和可解释的方法.
- 包括结构描述符进一步提高了预测准确性,证明了协同效应.
- 这种DMS驱动的方法为传统基于结构的模型提供了一个有希望的替代方案,特别是在处理较小的数据集时.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
65
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
65
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
81
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
81
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Outliers and Influential Points
4.1K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.1K


