使用解释性机器学习方法,对三种典型的中国大米品种的磨度进行预测建模
Liu Yang1, Zilong Xu1, Xuan Xiao1
1College of Mechanical Engineering, Wuhan Polytechnic University, Wuhan, China.
Journal of food science
|September 1, 2024
概括
准确预测米磨度 (DOM) 对于防止营养损失至关重要. 这项研究开发了一种基于图像的机器学习模型,在预测残留树脂层 (DOR) 的程度方面达到超过91%的准确性.
科学领域:
- 农业工程 农业工程
- 食品科学 食品科学 食品科学
- 计算机视觉 计算机视觉
背景情况:
- 过度磨导致严重的经济和营养损失.
- 精确检测和预测米磨度 (DOM) 是具有挑战性的,特别是在中等加工过程中.
- 保持大米的质量和营养需要精确控制磨粉过程.
研究的目的:
- 开发一种自动化,非破坏性和具有成本效益的方法来预测米加工质量.
- 建立一个可靠的模型来预测残留的树脂层 (DOR) 的程度及其与DOM的关系.
- 识别影响削质量预测的关键图像特征.
主要方法:
- 一个定制的谷物图像采集平台被用于捕获大米图像.
- 使用图像处理技术从米粒中提取颜色,纹理和形状特征.
- 机器学习模型,包括Catboost,使用交叉验证和网格搜索进行DOR预测,进行了开发和优化.
主要成果:
- 优化的Catboost模型实现了91.24%的预测准确度,精度,回忆和F1得分超过90%.
- 沙普利的附加解释显示,颜色特征 (例如,YCbCr-Cb_ske) 是最重要的,其次是纹理 (例如,GLCM-对比度) 和形状.
- 该模型在准确度方面显著改善,从最初的84.28%提高到优化的91.24%.
结论:
- 图像处理与机器学习相结合,为自动化和非破坏性的米度预测提供了可行的解决方案.
- 特性重要性分析为改进未来削质量预测模型提供了宝贵的见解.
- 这种方法可以指导改善米加工实践,以提高营养保留和减少破米产量.
更多相关视频
10:25Construction of Models for Nondestructive Prediction of Ingredient Contents in Blueberries by Near-infrared Spectroscopy Based on HPLC Measurements
Published on: June 28, 2016
10.6K
05:22Transverse Sectioning of Mature Rice Oryza sativa L. Kernels for Scanning Electron Microscopy Imaging Using Pipette Tips as Immobilization Support
Published on: January 25, 2022
3.6K
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Response Surface Methodology
98
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
98
Light Acquisition
8.4K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
8.4K
Mechanistic Models: Compartment Models in Individual and Population Analysis
33
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
33
