Repeated time-series cross-validation: A new method to improved COVID-19 forecast accuracy in Malaysia
Azlan Abdul Aziz1,2, Marina Yusoff3,4,5, Wan Fairos Wan Yaacob6
1College of Computing, Informatics and Mathematics, Universiti Teknologi MARA (UiTM) Cawangan Perlis, Arau 02600, Perlis, Malaysia.
Abstract:
Forecasting COVID-19 cases is challenging, and inaccurate forecast values will lead to poor decision-making by the authorities. Conversely, accurate forecasts aid Malaysian government authorities and agencies (National Security Council, Ministry of Health, Ministry of Finance, Ministry of Education, and Ministry of International Trade and Industry) and financial institutions in formulating action plans, regulations, and legal acts to control COVID-19 spread in the country. Therefore, this study proposes Repeated Time-Series Cross-Validation, a new data-splitting strategy to identify the best forecasting model that is capable of producing the lowest error measures value and a high percentage of forecast accuracy for COVID-19 prediction in Malaysia. Some of the highlights of the proposed method are:•A total of 21 models, five data partitioning sets, and four error measures to improve the forecast accuracy of daily COVID-19 cases in Malaysia.•The best model selected produces the lowest error measure value for the Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and Mean Absolute Scaled Error (MASE).•The average 8-day forecast accuracy is 90.2 %. The lowest and highest forecast accuracy was 83.7 % and 98.7 %.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Steps in Outbreak Investigation
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Statistical Methods for Analyzing Epidemiological Data
Improving Translational Accuracy


