一种简单且可扩展的内核密度方法,用于原子机器学习中可靠的不确定性量化
Daniel Willimetz1, Lukáš Grajciar1
1Department of Physical and Macromolecular Chemistry, Charles University, Hlavova 8, Praha 2, Prague 12800, Czech Republic.
The journal of physical chemistry letters
|October 16, 2025
概括
本研究介绍了一种GPU加速的不确定性量化框架,使用内核密度估计 (KDE) 来识别材料科学中机器学习模型的不可靠预测. 该方法通过检测数据缺口而确保模型可靠性,而不需要重新训练复杂的模型.
科学领域:
- 材料科学 材料科学 材料科学
- 计算化学的计算化学
- 机器学习 机器学习
背景情况:
- 机器学习 (ML) 模型对于预测材料特性和加速模拟至关重要.
- 模型可靠性取决于培训数据的代表性,在高维空间中存在挑战.
- 目前用于不确定性量化的方法可能是计算上昂贵的,通常需要模型合集.
研究的目的:
- 为材料科学中的ML模型开发一个可扩展和高效的不确定性量化 (UQ) 框架.
- 为评估预测可信度提供一种模型不可知度量.
- 为了实现实际部署,并提高ML模型在科学应用中的可解释性.
主要方法:
- 实现了一个GPU加速的UQ框架,使用k-最近邻近内核密度估计 (KDE).
- 在描述器空间中使用主要组件分析 (PCA) 来减少维度.
- 在各种化学系统,ML模型,描述器和材料特性中验证了框架.
主要成果:
- 基于KDE的不确定性得分有效地识别了样本稀少的区域和推断配置.
- UQ指标与传统的基于集团的不确定性指标有很强的相关性.
- 该框架成功地突出了ML模型预测不那么值得信赖的领域.
结论:
- 拟议的基于KDE的UQ框架为评估ML模型可靠性提供了一个实用和可转移的解决方案.
- 它增强了材料科学中的ML模型的可解释性和稳定性.
- 这种方法通过量化预测不确定性来促进ML模型的部署准备.
相关概念视频
Propagation of Uncertainty from Systematic Error
1.3K
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
1.3K
The Uncertainty Principle
31.3K
Werner Heisenberg considered the limits of how accurately one can measure properties of an electron or other microscopic particles. He determined that there is a fundamental limit to how accurately one can measure both a particle’s position and its momentum simultaneously. The more accurate the measurement of the momentum of a particle is known, the less accurate the position at that time is known and vice versa. This is what is now called the Heisenberg uncertainty principle. He...
31.3K
Uncertainty: Overview
1.5K
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
1.5K
The Quantum-Mechanical Model of an Atom
56.5K
Shortly after de Broglie published his ideas that the electron in a hydrogen atom could be better thought of as being a circular standing wave instead of a particle moving in quantized circular orbits, Erwin Schrödinger extended de Broglie’s work by deriving what is now known as the Schrödinger equation. When Schrödinger applied his equation to hydrogen-like atoms, he was able to reproduce Bohr’s expression for the energy and, thus, the Rydberg formula governing hydrogen spectra.
56.5K
Propagation of Uncertainty from Random Error
1.6K
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
1.6K
Uncertainty: Confidence Intervals
10.1K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
10.1K


