通过从机器学习模型中对不确定性的强有力的量化来提高分子性质预测的可靠性
Alex Kötter1, Kanishka Singh2, Hans Matter2
1Digital R&D Large Molecule Research, Sanofi-Aventis Deutschland GmbH, 65926 Frankfurt am Main, Germany.
Journal of chemical information and modeling
|October 8, 2025
概括
量化机器学习 (ML) 模型不确定性对于分子性质预测至关重要. 这项研究揭示了当前不确定性量化 (UQ) 方法的局限性,特别是复杂的结构-活动关系,并引入了一种强大的新的UQ方法.
科学领域:
- 计算化学是一种计算化学.
- 机器学习 机器学习
- 药物发现 药物发现
背景情况:
- 预测不确定性量化 (UQ) 对于机器学习 (ML) 在分子性质预测中至关重要.
- 现有的UQ方法在识别与化学空间复杂性和数据表示相关的错误方面面临挑战.
研究的目的:
- 在分子活动预测中分析错误源和UQ性能之间的关系.
- 评估数据分割策略对UQ方法评估的影响.
- 为分子ML模型开发一个改进的UQ方法.
主要方法:
- 对分子活动数据集的流行的UQ方法的分析.
- 调查错误来源:化学空间区域和训练数据表示.
- 开发和验证一种新的UQ方法.
主要成果:
- 几种UQ方法无法在的结构-活性关系 (SAR) 区域中检测出预测不佳的化合物.
- 数据分割策略显著影响了UQ的表现.
- 拟议的UQ方法在各种评估场景中显示出强有力的改进.
结论:
- 当前的UQ方法在确定分子性质预测中的特定错误源方面存在局限性.
- 一种新的,强大的UQ方法提高了ML引导分子发现的可靠性.
- 开发的UQ方法在积极学习环境中是有效的.
相关概念视频
Uncertainty: Overview
1.6K
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
1.6K
Propagation of Uncertainty from Systematic Error
1.3K
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
1.3K
Propagation of Uncertainty from Random Error
1.7K
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
1.7K
Uncertainty in Measurement: Accuracy and Precision
99.7K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
99.7K
Uncertainty: Confidence Intervals
10.1K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
10.1K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K


