Related Experiment Video
Updated: Oct 13, 2025

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
Forecasting COVID-19 outbreak progression using hybrid polynomial-Bayesian ridge regression model
1Mathematic and Computing Department, Indian Institute of Technology (Indian School of Mines), Dhanbad, Jharkhand India.
Insights
This study introduces a hybrid machine learning model for accurate COVID-19 case forecasting, incorporating uncertainty using Bayesian Ridge Regression and polynomial methods. The model effectively predicts future cases while managing prediction uncertainty.
Area of Science:
- Epidemiology
- Machine Learning
- Computational Statistics
Background:
- The COVID-19 pandemic, caused by SARS-CoV-2, presented unprecedented global health challenges.
- Accurate forecasting of infected cases is crucial for effective public health response and resource allocation.
- Existing forecasting methods often struggle with uncertainty and incorporating new data efficiently.
Purpose of the Study:
- To propose a novel hybrid machine learning model for predicting COVID-19 cases.
- To develop a model that accurately forecasts infected cases while quantifying prediction uncertainty.
- To create a flexible mathematical model capable of incorporating prior knowledge and new data.
Main Methods:
- A hybrid model combining Bayesian Ridge Regression with an n-degree Polynomial was developed.
- Probabilistic distributions were used for estimating the dependent variable, moving beyond traditional deterministic methods.
- L² (Ridge) Regularization was implemented to prevent model overfitting and enhance generalization.
- The model incorporates prior knowledge and posterior distributions for efficient data assimilation.
Main Results:
- The hybrid model demonstrated high accuracy in predicting COVID-19 cases.
- The model successfully quantified the uncertainty associated with its predictions.
- Case studies in the United States, Italy, and Spain validated the model's forecasting capabilities.
- Forecasts were based on publicly available data up to May 11, 2020.
Conclusions:
- The proposed hybrid model offers a robust approach to COVID-19 forecasting.
- The model's ability to handle uncertainty and incorporate new data makes it valuable for public health.
- Further research and evolution of the model are recommended for ongoing pandemic management.
Abstract:
In 2020, Coronavirus Disease 2019 (COVID-19), caused by the SARS-CoV-2 (Severe Acute Respiratory Syndrome Corona Virus 2) Coronavirus, unforeseen pandemic put humanity at big risk and health professionals are facing several kinds of problem due to rapid growth of confirmed cases. That is why some prediction methods are required to estimate the magnitude of infected cases and masses of studies on distinct methods of forecasting are represented so far. In this study, we proposed a hybrid machine learning model that is not only predicted with good accuracy but also takes care of uncertainty of predictions. The model is formulated using Bayesian Ridge Regression hybridized with an n-degree Polynomial and uses probabilistic distribution to estimate the value of the dependent variable instead of using traditional methods. This is a completely mathematical model in which we have successfully incorporated with prior knowledge and posterior distribution enables us to incorporate more upcoming data without storing previous data. Also, L2 (Ridge) Regularization is used to overcome the problem of overfitting. To justify our results, we have presented case studies of three countries, -the United States, Italy, and Spain. In each of the cases, we fitted the model and estimate the number of possible causes for the upcoming weeks. Our forecast in this study is based on the public datasets provided by John Hopkins University available until 11th May 2020. We are concluding with further evolution and scope of the proposed model.
Related Concept Videos
Steps in Outbreak Investigation
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Statistical Methods for Analyzing Epidemiological Data
Causality in Epidemiology

