Dengue in Tomorrow: Predictive Insights From ARIMA and SARIMA Models in Bangladesh: A Time Series Analysis
Pratyay Hasan1, Tazdin Delwar Khan2, Ishteaque Alam3
1Department of Respiratory Medicine (OPD) Dhaka Medical College Hospital Dhaka Bangladesh.
Backgrounds And Aims:
Dengue fever has been a continued public health problem in Bangladesh, with a recent surge in cases. The aim of this study was to train ARIMA and SARIMA models for time series analysis on the monthly prevalence of dengue in Bangladesh and to use these models to forecast the dengue prevalence for the next 12 months.
Methods:
This secondary data-based study utilizes AutoRegressive Integrated Moving Average (ARIMA) and Seasonal AutoRegressive Integrated Moving Average (SARIMA) models to forecast dengue prevalence in Bangladesh. Data was sourced from the Institute of Epidemiology Disease Control and Research (IEDCR) and the Directorate General of Health Services (DGHS). STROBE Guideline for observational studies was followed for reporting this study.
Results:
The ARIMA (1,1,1) and SARIMA (1,1,2) models were identified as the best-performing models. The forecasts indicate a steady dengue prevalence for 2024 according to ARIMA, while SARIMA predicts significant fluctuations. It was observed that ARIMA (1,1,1) and SARIMA (1,2,2) (1,1,2) were the most suitable models for prediction of dengue prevalence.
Conclusion:
These models offer valuable insights for healthcare planning and resource allocation, although external factors and complex interactions must be considered. Dengue prevalence is expected to rise in future in Bangladesh.
Related Concept Videos
Steps in Outbreak Investigation
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...


