使用信息理论在原子化机器学习中对完整性,不确定性和异常值的无模型估计
Daniel Schwalbe-Koda1,2, Sebastien Hamel3, Babak Sadigh3
1Lawrence Livermore National Laboratory, Livermore, CA, 94550, USA. dskoda@ucla.edu.
Nature communications
|April 29, 2025
概括
我们开发了一个无模型的框架,用信息来量化原子模拟中的信息. 这种方法增强了机器学习潜力的发展,并使模拟可靠的不确定性量化.
科学领域:
- 原子化的机器学习.
- 计算材料科学 计算材料科学
- 数据驱动的建模.
背景情况:
- 原子式机器学习 (ML) 通常使用无监督学习或模型预测进行数据分析.
- 准确的信息描述对于训练集,不确定性量化 (UQ) 和提取物理见解至关重要.
研究的目的:
- 引入一个严格的,无模型的理论框架来量化原子模拟中的信息内容.
- 为数据驱动的原子模型提供一个通用的工具.
主要方法:
- 使用原子中心环境的信息量化信息内容.
- 在这个信息框架的基础上开发一种无模型的UQ方法.
主要成果:
- 信息解释了ML潜在发展中的启发式,包括训练集大小和数据集最佳性.
- 拟议的UQ方法可靠地预测认识体系的不确定性.
- 该方法有效地检测到分布外的样本和罕见事件,如核化.
结论:
- 开发的框架提供了一个强大的,无模型的工具,用于在原子模拟中分析信息.
- 这种方法集成了机器学习,模拟和物理可解释性,用于增强数据驱动的建模.
相关概念视频
Propagation of Uncertainty from Systematic Error
345
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
345
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
25
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
25
Propagation of Uncertainty from Random Error
500
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
500
Uncertainty: Overview
387
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
387
Mechanistic Models: Compartment Models in Individual and Population Analysis
13
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
13
Uncertainty: Confidence Intervals
2.9K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
2.9K


