通过选择包含最合适样本的子集来减少校准集,这些样本显示出最佳的预测能力
Jan P M Andries1, Yvan Vander Heyden2
1Research Group Analysis Techniques in the Life Sciences, Avans Hogeschool, University of Professional Education, P.O. Box 90116, 4800 RA, Breda, the Netherlands.
在近红外 (NIR) 光谱中选择校准样本的新方法,称为最佳预测校准子集 (OPCS),在保持预测能力的同时显著减少数据集大小. 这种方法确保不包括异常值,提供更有效的建模策略.
科学领域:
- 分析化学 分析化学
- 频谱学是一种光谱学.
- 化学测量 化学测量 化学测量
背景情况:
- 近红外 (NIR) 光谱是一种快速,非侵入性和具有成本效益的分析技术.
- 在NIR光谱学中经常使用大型校准集,但可以在不丢失信息的情况下选择有信息的子集.
- 有效的样本子集选择对于开发强大且具有成本效益的NIR模型至关重要.
研究的目的:
- 提出和评估一种新的方法来选择NIR光谱中的最佳校准样本子集.
- 开发一种方法,减少校准集大小,同时保持或增强预测能力.
- 为改进NIR模型开发引入最佳预测校准子集 (OPCS) 方法.
主要方法:
- 开发了一种新的样本子集选择方法,即最佳预测校准子集 (OPCS).
- 该方法使用全局部分最小平方 (PLS) 模型和交叉模型验证 (CMV) 来确定最适合的样本.
- 来自重复双十字验证的"一个标准错误规则"在CMV内应用,以最佳地确定PLS复杂性.
主要成果:
- 与肯纳德-斯通方法和共同指导方针相比,OPCS方法的结果是统计学上显著较小的样本子集.
- 基于OPCS的模型展示了与使用完整原始校准集构建的模型相比的预测能力.
- 由于OPCS方法本质上排除了异常值,因为只选择最合适的校准样本.
结论:
- 在NIR光谱学中,OPCS方法提供了一种有效的策略,可以在不影响预测性能的情况下减少校准集大小.
- 这种方法比传统方法具有优势,因为它确保了模型的子集代表性,而不是整个数据集,并排除了异常值.
- 用CMV进行样本选择和"一个标准误差规则"的整合代表了化学计量建模中的重大创新.
更多相关视频
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
10:25Construction of Models for Nondestructive Prediction of Ingredient Contents in Blueberries by Near-infrared Spectroscopy Based on HPLC Measurements
Published on: June 28, 2016
相关概念视频
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
Calibration Curves: Correlation Coefficient
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Instrument Calibration
Analytical Balance Calibration
An analytical balance measures mass and requires regular calibration to...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
