Related Experiment Video
Updated: Jan 14, 2026

High-Throughput, In-Field Screening of Photosynthetic Efficiency in Crop Plants Using an Autonomous Robot
Published on: January 9, 2026
An interpretable machine learning approach based on SHAP, Sobol and LIME values for precise estimation of daily
Ahmed Elbeltagi1,2, Aman Srivastava3, Xinchun Cao4,5
1College of Agricultural Science and Engineering, Hohai University, Nanjing, 211100, China.
Abstract:
Increasing water scarcity and climate variability have intensified the need for precise agricultural irrigation management. Accurate estimation of crop coefficients (Kc) is critical for determining crop water requirements, especially in arid and semi-arid regions. However, conventional methods for estimating Kc often rely on generalized plant characteristics, which may not account for local climatic variations. In this study, we address this challenge by predicting the daily crop coefficient for soybean using four machine learning models: Extreme Gradient Boosting (XGBoost), Extra Tree (ET), Random Forest (RF), and CatBoost. These models were trained on meteorological data from Suhaj Governorate, Egypt, spanning 1979-2014. Additionally, SHapley Additive exPlanations (SHAP), Sobol sensitivity analysis, and Local Interpretable Model-agnostic Explanations (LIME) were applied to evaluate model interpretability and consistency with physical processes. Among the models evaluated, the ET model achieved the highest accuracy, with r = 0.96, NSE = 0.93, RMSE = 0.05, and MAE = 0.02. XGBoost and RF also performed well, each obtaining r = 0.96, NSE = 0.92, RMSE = 0.06, and MAE = 0.02. In comparison, CatBoost demonstrated slightly lower accuracy, with r = 0.95, NSE = 0.91, RMSE = 0.06, and MAE = 0.02. SHAP and Sobol analyses consistently identified the antecedent crop coefficient [[Formula: see text]] and solar radiation (Sin) as the most influential variables. LIME results revealed localized variations in predictions, reflecting dynamic crop-climate interactions. This study underscores the importance of integrating interpretable machine learning models to enhance both predictive accuracy and reliability while maintaining alignment with critical physical processes. The proposed framework offers a robust tool for improving daily Kc estimation, thereby supporting more sustainable irrigation practices and climate-resilient agriculture.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
Calculating and Interpreting the Linear Correlation Coefficient
Light Acquisition
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

