基于机器学习的全球密度依赖范围分离参数的准确预测.
Corentin Villot1, Tong Huang1, Ka Un Lao1
1Department of Chemistry, Virginia Commonwealth University, Richmond, Virginia 23284, USA.
The Journal of chemical physics
|July 24, 2023
概括
一个新的XGBoost模型准确地预测了长距离校正函数 (LRC-ωPBE) 的范围分离参数,绕过了复杂的计算. 这种机器学习方法可以对各种化学系统进行快速,高效的预测.
科学领域:
- 计算化学计算化学
- 机器学习在量子力学中的应用
- 密度函数理论 密度函数理论
背景情况:
- 准确预测电子特性需要精确的密度函数近似值.
- 长距离校正 (LRC) 函数,如LRC-ωPBE,提高准确性,但通常需要系统特定的参数调整.
- 确定最佳范围分离参数 (ω) 对LRC的功能性能至关重要.
研究的目的:
- 开发一个准确和高效的机器学习模型来预测LRC-ωPBE的全球密度依赖范围分离参数 (ωGDD).
- 仅使用原子坐标来实现快速,计算上便宜的 ωGDD 预测.
- 证明开发模型在预测分子性质和相互作用方面的实用性.
主要方法:
- 开发一个XGBoost机器学习模型 (ωGDDML) 使用本地原子环境的指纹和距离直方图.
- 该模型的培训和验证是在11466种不同的化学系统的大数据集上进行的.
- 应用LRC-ωPBE ((ωGDDML) 函数来预测极化性和非共价相互作用.
主要成果:
- ωGDDML模型在7046个复合体的测试集上实现了高精度,平均绝对误差为0.001117a0-1.
- 观察到很好的可转移性,只有0.07%的系统显示的错误大于0.01 a0-1.
- LRC-ωPBE ((ωGDDML) 在预测极化性方面表现出卓越的性能,并绕过了对非共价相互作用的传统初始系统特定调整的需求.
结论:
- 开发的 ωGDDML 模型提供了一个准确,高效和可转移的方法来确定 LRC-ωPBE 的范围分离参数.
- 这种数据驱动的方法通过消除对参数预测进行电子结构计算的需求,显著降低了计算成本.
- 物理启发的LRC-ωPBE与数据驱动的 ωGDDML模型的融合为量子力学计算提供了协同效益.
更多相关视频
12:26Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM
Published on: October 11, 2016
13.4K
08:47Author Spotlight: UAV Remote Sensing for Efficient Invasive Plant Biomass Estimation
Published on: February 9, 2024
1.5K
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
64
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
64
Habitat Fragmentation
17.7K
Habitat fragmentation describes the division of a more extensive, continuous habitat into smaller, discontinuous areas. Human activities such as land conversion, as well as slower geological processes leading to changes in the physical environment, are the two leading causes of habitat fragmentation. The fragmentation process typically follows the same steps: perforation, dissection, fragmentation, shrinkage, and attrition.
17.7K
Range
11.7K
The range is one of the measures of variation. It can be defined as the difference between a dataset's highest and lowest values. For example, in the study of seven 16-ounce soda cans, the filled volume of soda was measured, thus producing the following amount (in ounces) of soda:
15.9; 16.1; 15.2; 14.8; 15.8; 15.9; 16.0; 15.5
Measurements of the amount of soda in a 16-ounce can vary since different subjects record these measurements or since the exact amount - 16 ounces of liquid, was not...
15.9; 16.1; 15.2; 14.8; 15.8; 15.9; 16.0; 15.5
Measurements of the amount of soda in a 16-ounce can vary since different subjects record these measurements or since the exact amount - 16 ounces of liquid, was not...
11.7K
Distribution and Dispersion
21.9K
To understand intra-specific interactions in populations, scientists measure the spatial arrangement of species individuals. This geographic arrangement is known as the species distribution or dispersion. Highly territorial species exhibit a uniform distribution pattern, in which individuals are spaced at relatively equal distances from one another. Species that are highly tied to particular resources, such as food or shelter, tend to concentrate around those resources, and thus exhibit a...
21.9K
Hybrid Zones
17.1K
Hybrid zones are narrow regions where two closely related species interact, mate, and produce hybrids. Relative to either parent species, hybrids may possess distinct phenotypic or genetic differences that impact their survival and reproductive success. The genetic variances introduced by hybridization influence species diversity and speciation processes within the hybrid zone.
17.1K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
