预测颗粒耐用性指数在商业料厂使用多重线性回归与变量选择和缩小维度的预测
Jihao You1, Dan Tulpan1, Cheryl Krziyzek2
1Department of Animal Biosciences, University of Guelph, Guelph, ON, Canada.
Journal of animal science
|February 4, 2025
概括
预测颗粒质量对于料制造至关重要. 这项研究开发了统计模型来预测颗粒耐用性指数 (PDI),确定影响料质量的温度和脂肪含量等关键变量.
科学领域:
- 农业工程 农业工程
- 动物科学动物科学
- 统计建模 统计建模
背景情况:
- 用颗粒耐用性指数 (PDI) 衡量颗粒质量,对于料制造效率和动物营养至关重要.
- 由于许多复杂的工艺变量,控制颗粒质量是一项挑战.
- 现有的预测方法使用有限的实证模型或复杂的机器学习方法.
研究的目的:
- 开发统计回归模型,用于预测商业料制造业的PDI.
- 识别和描述颗粒质量与55个潜在影响变量之间的关系.
主要方法:
- 从一家商业料厂收集了2691个观察的数据集.
- 使用提升正常性 (tPDI) 的Box-Cox方法转换了PDI变量.
- 开发了三种多重回归模型,使用前向选择,主要组件分析和部分最小方程,然后选择变量.
主要成果:
- 前向选择模型 (模型1) 使用9个变量,显示出比PCA和PLS模型更优异的性能.
- 模型1在训练和测试数据上显示出一致的预测准确性,具有低误差指标 (MAE,RMSPE) 和良好的一致性相关系数.
- 在模型1中,扩展温度,脂肪含量,ADF含量和室内湿度 (颗粒剂) 被确定为影响TPDI的最有影响力的变量.
结论:
- 统计回归模型,特别是前选择方法,可以有效地预测颗粒质量 (PDI).
- 温度,脂肪和ADF含量等关键变量显著影响颗粒的耐用性.
- 这些模型为料厂提供了有价值的工具,可以预测和理解影响颗粒质量的因素.
更多相关视频
06:50O-cresol Concentration Online Measurement Based On Near Infrared Spectroscopy Via Partial Least Square Regression
Published on: November 8, 2019
6.5K
10:25Construction of Models for Nondestructive Prediction of Ingredient Contents in Blueberries by Near-infrared Spectroscopy Based on HPLC Measurements
Published on: June 28, 2016
10.6K
相关概念视频
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Residual Plots
4.5K
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
4.5K
Variation
6.7K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.7K
Microsoft Excel: Regression Analysis
442
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
442
