强化采样稳健的分子数据集与基于不确定性的集体变量
Aik Rui Tan1, Johannes C B Dietschreit1,2, Rafael Gómez-Bombarelli1
1Department of Materials Science and Engineering, Massachusetts Institute of Technology, Cambridge, Massachusetts 02139, USA.
The Journal of chemical physics
|January 15, 2025
概括
本研究引入了一种使用模型不确定性的新方法,用于指导机器学习潜力的数据生成. 这种方法有效地探索复杂的分子配置,提高模型的准确性和稳定性.
科学领域:
- 计算化学是一种计算化学.
- 材料科学是一种材料科学.
- 机器学习 机器学习
背景情况:
- 生成代表性数据集对于强大的机器学习原子间潜力 (MLIP) 是至关重要的.
- 复杂的分子系统由于复杂的潜在能量表面和众多的局部最小值带来了挑战.
- 传统的数据生成方法通常是低效的,或者无法捕获关键配置.
研究的目的:
- 开发一个数据效率高的积极学习策略,用于为MLIPs生成高质量的数据集.
- 为了应对探索复杂的潜在能量表面和克服能源障碍的挑战.
- 为了提高机器学习的原子间潜力的稳定性和准确性.
主要方法:
- 利用模型不确定性作为一个集体变量 (CV) 来指导数据采集.
- 从单个模型中采用高斯混合模型的不确定性度量.
- 利用偏向的分子动力学模拟来对配置空间进行有针对性的探索.
主要成果:
- 在克服能源障碍和探索前所未有的能源最小值方面表现出有效性.
- 在积极学习框架中成功增强数据集.
- 验证了对二和散装二氧化系统的方法.
结论:
- 提出的以不确定性为导向的方法显著提高了MLIPs的数据生成效率.
- 这种方法可以通过专注于化学相关,不确定的区域来实现更强大,更准确的分子建模.
- 积极学习框架可以通过对目标数据采集的不确定性量化来有效增强.
相关概念视频
Uncertainty: Confidence Intervals
3.1K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
3.1K
Uncertainty: Overview
509
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
509
Propagation of Uncertainty from Systematic Error
462
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
462
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
Propagation of Uncertainty from Random Error
636
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
636
Distributions to Estimate Population Parameter
4.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.0K


