Related Experiment Video
Updated: Nov 4, 2025

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
Machine learning models for the prediction of the SEIRD variables for the COVID-19 pandemic based on a deep
Yullis Quintero1, Douglas Ardila1, Edgar Camargo2
1GIDITIC, Universidad EAFIT, Colombia.
Abstract:
The SEIRD (Susceptible, Exposed, Infected, Recovered, and Dead) model is a mathematical model based on dynamic equations; widely used for characterization of the COVID-19 pandemic. In this paper, a different approach has been discussed, which is the development of predictive models for the SEIRD variables that have been based on the historical data collected, and the context variables to where this model has been applied to. Particularly, the context variables examined in this paper include total population, number of people over 65 years old, poverty index, morbidity rates, average age, and population density. For the construction of the SEIRD predictive models, this study encompasses a deep analysis of the dependence of these variables and also, their relationship with the context variables. Hence, before the development of predictive models using machine learning techniques, a methodology to analyze the interdependence of the SEIRD variables has been proposed. The dependence with the context variables is also discussed; to avoid the curse of dimensionality and multicollinearity problems, leading to better results and the reduction of the computational cost. Finally, several prediction models based on varied machine learning techniques and inputs are considered, these include temporal interdependence, temporal intra-dependence, and dependence with context variables. Each of the predictive models has been studied, as well as their quality of prediction. This paper focuses on the analysis of the quality of this approach, applied in Colombia, obtaining the results about the performance of the predictive models for the SEIRD variables. The results are very encouraging since the values obtained with the quality metrics are quite good for different prediction horizons.
Related Concept Videos
Steps in Outbreak Investigation
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Statistical Methods for Analyzing Epidemiological Data
Causality in Epidemiology
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

