一个分层的RF-XGBoost模型用于短期周期农产品销售预测
Jiawen Li1,2, Binfan Lin1, Peixian Wang1
1School of Computer Science, Guangdong Polytechnic Normal University, Guangzhou 510665, China.
Foods (Basel, Switzerland)
|September 28, 2024
概括
本研究介绍了一个层次的RF-XGBoost模型,用于准确预测农业销售,通过改进需求预测来显著减少食物浪费. 该模型提高了从农场到餐桌的供应链效率.
科学领域:
- 农业经济学 农业经济学
- 数据科学数据科学数据科学
- 供应链管理 供应链管理
背景情况:
- 农产品短周期销售预测对于通过调整供应与需求来最大限度地减少食物浪费至关重要.
- 由于不确定的因素,由于波动和不连续的销售数据,预测准确性受到挑战.
- 现有的模型经常与农业市场动态固有的复杂性作斗争.
研究的目的:
- 开发和评估一种新的等级预测模型,以提高农产品短期销售预测.
- 提高需求预测的准确性,从而减少食物浪费和优化农业供应链.
- 通过先进的建模方法来解决销售数据的波动性和不连续性.
主要方法:
- 开发了一种将随机森林 (RF) 和极端梯度提升 (XGBoost) 结合在一起的等级模型.
- 第一个层使用射频与灰色关系分析 (GRA) 进行初始预测和残余提取.
- 第二层使用XGBoost在剩余集群特征上进行精细预测.
主要成果:
- 拟议的RF-XGBoost模型与独立的RF和XGBoost相比表现出更高的性能.
- 与RF和XGBoost相比,平均绝对百分比误差 (MAPE) 分别减少了10%和12%.
- 确定系数 (R2) 增加了22%和24%,表明预测准确度更高.
结论:
- 层次化的RF-XGBoost模型显著提高了短期农业销售预测的精度.
- 该模型的有效性在各种农产品中得到了验证,证明了其广泛适用性.
- 实施为优化供应链和减少食物浪费提供了巨大的好处.
更多相关视频
10:25Construction of Models for Nondestructive Prediction of Ingredient Contents in Blueberries by Near-infrared Spectroscopy Based on HPLC Measurements
Published on: June 28, 2016
10.6K
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
653
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Regression Analysis
5.6K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.6K
Microsoft Excel: Regression Analysis
504
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
504
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
Pharmacokinetic Models: Comparison and Selection Criterion
48
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
48
