在材料科学中通过近邻和代预测计算缺失数据的归算
Chunhui Xie1, Rui Li1, Yunqi Li1
1Department of Polymer Materials and Engineering, College of Materials and Metallurgy, Guizhou University, Guiyang 550025, P. R. China.
Journal of chemical theory and computation
|December 26, 2024
概括
MatImpute有效地解决了材料科学数据集中缺失的数据. 这种新的归算策略将最大限度地减少错误,并提高机器学习应用程序的数据质量.
科学领域:
- 材料科学 材料科学 材料科学
- 数据科学数据科学数据科学
- 机器学习 机器学习
背景情况:
- 数据缺失是材料科学中常见的挑战,影响了统计分析,大数据和机器学习.
- 现有的归算策略缺乏在材料科学背景下对可靠性的严格评估.
研究的目的:
- 为了对材料科学数据集的六种数据归算策略的性能进行基准和评估.
- 引入和验证一种新的归算方法,MatImpute,以提高数据质量.
主要方法:
- 在七个材料科学数据集中比较了Mean,MissForest,HyperImpute,Gain,Sinkhorn和MatImpute.
- 通过使用根平均平方误差 (RMSE),瓦斯斯坦距离 (WD) 和数据集相关性收 (DCC) 评估的归算诱导错误 (IIEs).
- 在不同缺失数据比率和类型 (MAR,MCAR,MNAR) 和回归/分类任务中评估性能.
主要成果:
- MatImpute表现出卓越的性能,实现了最低的RMSE和WD,以及最高的DCC.
- 由于缺失数据比较高,推算诱导的错误增加,并遵循趋势 MAR < MCAR ≤ MNAR.
- 在机器学习预测模型中,MatImpute保持了最高的数据恢复保真度.
结论:
- 在材料科学中,MatImpute是一个非常有效的策略,用于归纳缺失的数据,其性能优于现有的方法.
- 该研究为评估归算策略提供了一个框架,并强调了数据质量对于可靠的材料科学研究的重要性.
- 发布MatImpute代码以支持创建高质量的材料科学数据集.
相关概念视频
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Predicting Molecular Geometry
34.1K
VSEPR Theory for Determination of Electron Pair Geometries
34.1K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
40
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
40
Predicting Products: Substitution vs. Elimination
11.4K
When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other:
The following factors can influence the mechanisms competing against each other:
11.4K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Mechanistic Models: Compartment Models in Individual and Population Analysis
27
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
27


