Machine learning ensembles, neural network, hybrid and sparse regression approaches for weather based rainfed cotton
Girish R Kashyap1, Shankarappa Sridhara2, Konapura Nagaraja Manoj1
1Centre for Climate Resilient Agriculture, Keladi Shivappa Nayaka University of Agricultural and Horticultural Sciences, Shivamogga, Karnataka, 577204, India.
Abstract:
Cotton is a major economic crop predominantly cultivated under rainfed situations. The accurate prediction of cotton yield invariably helps farmers, industries, and policy makers. The final cotton yield is mostly determined by the weather patterns that prevail during the crop growing phase. Crop yield prediction with greater accuracy is possible due to the development of innovative technologies which analyses the bigdata with its high-performance computing abilities. Machine learning technologies can make yield prediction reasonable and faster and with greater flexibility than process based complex crop simulation models. The present study demonstrates the usability of ML algorithms for yield forecasting and facilitates the comparison of different models. The cotton yield was simulated by employing the weekly weather indices as inputs and the model performance was assessed by nRMSE, MAPE and EF values. Results show that stacked generalised ensemble model and artificial neural networks predicted the cotton yield with lower nRMSE, MAPE and higher efficiency compared to other models. Variable importance studies in LASSO and ENET model found minimum temperature and relative humidity as the main determinates of cotton yield in all districts. The models were ranked based these performance metrics in the order of Stacked generalised ensemble > ANN > PCA ANN > SMLR ANN > LASSO> ENET > SVM > PCA SMLR > SMLR SVM > SMLR. This study shows that stacked generalised ensembling and ANN method can be used for reliable yield forecasting at district or county level and helps stakeholders in timely decision-making.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Precipitation and Co-precipitation
Precipitation Processes
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Multi-input and Multi-variable systems
In the absence...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:


