Large-scale multivariate forecasting models for Dengue - LSTM versus random forest regression
Elisa Mussumeci1, Flávio Codeço Coelho1
1Fundacao Getulio Vargas, Brazil.
Abstract:
Effective management of seasonal diseases such as dengue fever depends on timely deployment of control measures prior to the high transmission season. As the epidemic season fluctuates from year to year, the availability of accurate forecasts of incidence can be decisive in attaining control of such diseases. Obtaining such forecasts from classical time series models has proven a difficult task. Here we propose and compare machine learning models incorporating feature selection,such as LASSO and Random Forest regression with LSTM a deep recurrent neural network, to forecast weekly dengue incidence in 790 cities in Brazil. We use multivariate time-series as predictors and also utilize time series from similar cities to capture the spatial component of disease transmission. The LSTM recurrent neural network model attained the highest performance in predicting future incidence on dengue in cities of different sizes.
More Related Videos
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
Related Concept Videos
Steps in Outbreak Investigation
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Statistical Methods for Analyzing Epidemiological Data
Survival Tree
Building a Survival Tree
Constructing a...
