Related Experiment Videos
Year-independent prediction of rice grain protein content using machine learning with agronomy-aligned multi-year
Hyun-Jin Jung1, Yun-Ho Lee2, ChungGen Lee3
1Winter Crop Research Division, National Institute of Crop and Food Science, Rural Development Administration, Wanju, Republic of Korea.
Abstract:
Rice grain protein content exhibits substantial inter-annual variability due to interactions between management practices and environmental conditions, making robust prediction under realistic field conditions challenging. This study evaluated the year-independent predictability of rice grain protein content using multi-year field data and machine learning approaches, with a focus on data structuring and validation strategies. Models were trained using both plot-level and replicate-mean datasets, and evaluated using leave-one-year-out cross-validation to simulate prediction under unseen growing seasons. Replicate aggregation consistently improved predictive robustness by reducing plot-level variability, while year-independent validation provided a more conservative and realistic assessment of model generalization. Despite this stringent framework, predictive performance remained relatively stable, and residual diagnostics indicated no clear systematic bias across protein levels or nitrogen application rates. Anomaly-based analysis further suggested that deviations from year-variety baselines could be predicted with meaningful accuracy, indicating that relative variations beyond dominant temporal effects were effectively captured. These results suggest that aligning data structuring and validation strategies with agronomic experimental design may be important for reliable year-independent prediction. The proposed framework may provide a transferable approach for integrating machine learning with agronomic knowledge to support crop quality prediction and adaptive nitrogen management under variable field conditions.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Light Acquisition