Related Experiment Video
Updated: Oct 10, 2025

Cereal Crop Ear Counting in Field Conditions Using Zenithal RGB Images
Published on: February 2, 2019
Machine learning algorithms for soybean yield forecasting in the Brazilian Cerrado
Valter Barbosa Dos Santos1, Aline Moreno Ferreira Dos Santos1, José Reinaldo da Silva Cabral de Moraes2
1Graduate Program in Agronomy (Soil Science), State University of Sao Paulo (FCAV/UNESP), Jaboticabal, Brazil.
Machine learning models can predict soybean yield in Brazil's Matopiba region up to a month in advance. The random forest (RF) model demonstrated the highest accuracy, offering valuable insights for agricultural planning.
Area of Science:
- Agricultural Science
- Data Science
- Environmental Science
Background:
- Soybean productivity prediction is crucial for agricultural planning in Brazil's Matopiba region.
- Meteorological data from NASA-POWER and yield data from SIDRA/IBGE (2008-2017) were utilized.
- Several machine learning models were assessed for their predictive capabilities.
Purpose of the Study:
- To evaluate and compare various machine learning models for predicting soybean productivity.
- To forecast soybean yield up to one month in advance in the Matopiba agricultural frontier.
- To identify the most effective model for enhancing agricultural decision-making.
Main Methods:
- The study employed five machine learning models: random forest (RF), artificial neural networks, support vector machines with radial basis function (SVM_RBF), linear model, and polynomial regression.
- Cross-validation was used to assess model performance.
- Key performance metrics included R-squared (R²) for precision, root mean square error (RMSE) for accuracy, and mean error of the estimate (EME) for trend.
Main Results:
- The random forest (RF) algorithm exhibited the highest performance, achieving an R² of 0.81, an RMSE of 176.93 kg/ha, and an EME of 1.99 kg/ha.
- The SVM_RBF model showed the lowest performance with an R² of 0.74, an RMSE of 213.58 kg/ha, and an EME of -15.06 kg/ha.
- Predicted average yields by the models fell within the expected range, considering the region's historical average of 2,730 kg/ha.
Conclusions:
- All evaluated machine learning models demonstrated acceptable precision, accuracy, and trend indices for soybean yield prediction.
- These models can be applied to soybean crop yield prediction, considering regional specificities.
- The use of these algorithms serves as a valuable tool for agricultural planning and decision-making in Brazilian Cerrado soy-producing areas.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Light Acquisition
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Plant Breeding and Biotechnology
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...

