一种基于SHAP,Sobol和LIME值的可解释机器学习方法,用于精确估计每日大豆作物系数
Ahmed Elbeltagi1,2, Aman Srivastava3, Xinchun Cao4,5
1College of Agricultural Science and Engineering, Hohai University, Nanjing, 211100, China.
Scientific reports
|October 21, 2025
概括
机器学习模型准确地预测了用于大豆灌管理的每日作物系数 (Kc). 额外树模型显示了最高的准确性,提高了干旱地区的用水效率.
科学领域:
- 农业科学 农业科学
- 环境科学 环境科学
- 数据科学数据科学数据科学
背景情况:
- 越来越多的水资源短缺和气候变化需要精确的农业灌管理.
- 准确估计作物系数 (Kc) 对于确定作物用水需求至关重要,特别是在干旱和半干旱地区.
- 传统的Kc估计方法可能无法捕捉到当地的气候变化.
研究的目的:
- 使用机器学习模型预测大豆的每日作物系数 (Kc).
- 评估这些模型的可解释性和物理一致性.
- 为改善Kc估计和支持可持续灌提供一个强大的框架.
主要方法:
- 他们使用了四种机器学习模型 (Extra Tree,XGBoost,Random Forest,CatBoost).
- 模型是根据埃及苏哈吉省 (1979-2014) 的气象数据进行训练的.
- 使用夏普利添加式解释 (SHAP),索博尔灵敏度分析和局部可解释模型不可知解释 (LIME) 来评估可解释性.
主要成果:
- 额外树 (ET) 模型获得了最高的准确性 (r=0.96,NSE=0.93,RMSE=0.05,MAE=0.02).
- XGBoost和随机森林模型也表现出高性能.
- 前作物系数和太阳辐射被确定为SHAP和Sobol分析的关键影响变量.
结论:
- 可解释的机器学习模型提高了每日Kc估计的准确性和可靠性.
- 该研究强调了将模型预测与物理过程对齐对于稳健的农业管理的重要性.
- 拟议的框架支持可持续的灌实践和适应气候变化的农业.
相关概念视频
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Calibration Curves: Linear Least Squares
4.1K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
4.1K
Calculating and Interpreting the Linear Correlation Coefficient
7.9K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
7.9K
Light Acquisition
9.4K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
9.4K
Regression Analysis
8.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.0K
Residuals and Least-Squares Property
9.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.1K


