黑森数据是否提高了机器学习潜力的性能?
Austin Rodriguez1, Justin S Smith2, Jose L Mendoza-Cortes1,3
1Department of Chemical Engineering & Materials Science, Michigan State University, East Lansing, Michigan 48824, United States.
Journal of chemical theory and computation
|July 2, 2025
概括
将赫森矩阵训练集成到机器学习原子间潜能 (MLIP) 中,可以增强对新分子系统的推断. 这提高了反应建模和振动分析,尽管它增加了计算成本.
科学领域:
- 计算化学是一种计算化学.
- 材料科学是一种材料科学.
- 药物发现 药物发现
背景情况:
- 机器学习的原子间潜力 (MLIPs) 可以用量子精度预测能量和力.
- 在MLIP训练中使用强力改进了潜在能量表面预测.
- 黑森矩阵训练编码了关于潜在能量表面 (PES) 曲率的二级信息.
研究的目的:
- 评估黑森矩阵培训在MLIPs中的整合.
- 评估赫森训练对推断到不平衡几何学的影响.
- 分析黑森集成在MLIP中的好处和局限性,用于各种计算化学应用.
主要方法:
- 训练MLIP使用不同组合的能量,力和黑森数据.
- 在平衡和第一阶点几何学上评估模型性能.
- 测试使用小分子反应数据集对非平衡几何学的取值能力.
主要成果:
- 赫森训练有素的MLIP证明了对未见的分子系统的提取有所改善.
- 黑斯集成提高了反应路径建模和振动光谱预测的准确性.
- 尽管计算费用增加了,但Hessian培训可以减少有效MLIP模型所需的总数据.
结论:
- 在MLIP中,Hessian集成为预测分子性质和反应动态提供了显著的优势.
- 实践者应该权衡提高准确性和数据效率的好处与增加的计算成本.
- 这项工作提供了关于在计算化学研究中雇用赫森训练有素的MLIP的明智决策的见解.
相关概念视频
Survival Tree
157
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
157
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
207
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
207
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
100
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
100
Regression Analysis
6.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.0K
Residuals and Least-Squares Property
7.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.8K


