Related Experiment Video
Updated: Sep 15, 2025

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Machine Learning and Probabilistic Approaches for Forecasting COVID-19 Transmission and Cases
Md Sakhawat Hossain1,2, Ravi Goyal3, Natasha K Martin3
1Department of Public Health Sciences, Clemson University, Clemson, SC, USA.
Abstract:
Forecasting the effective reproductive number ( ) and COVID-19 case counts are critical for guiding public health responses. We developed a machine learning and probabilistic forecasting framework to predict and daily case counts at the county level in South Carolina (SC). Our approach utilized initial estimates from EpiNow2 R package refined with spatial (covariate-adjusted) smoothing. We then generated forecasts using an ensemble of regression, Random Forest, and XGBoost models, and predicted case counts with a probabilistic Poisson model. This ensemble-based approach consistently outperformed EpiNow2 across different forecast horizons (7-day, 14-day, and 21-day). In the first forecast period (November 11, 2020 - February 02, 2021), the ensemble achieved a median percentage agreement (PA) across counties of 94.4% (IQR: 93.8% - 95.3%) for 7-day ahead forecast, compared to 87.0% (IQR: 84.4% - 89.4%) from EpiNow2. In the second period (December 11, 2022 - March 04, 2023), the ensemble attained a 93.0% median PA across counties for Rt forecast (IQR: 91.3% - 94.1%), while EpiNow2 reached 86.8% (IQR: 82.5% - 89.2%). Similar trends were observed for case forecast, with the ensemble model demonstrating improved stability and performance. Combining spatial smoothing with ensemble modeling improves epidemic forecasting by enhancing predictive performance and robustness.
More Related Videos
10:46A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Steps in Outbreak Investigation
Statistical Methods for Analyzing Epidemiological Data
Causality in Epidemiology
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Principles of Disease Surveillance
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.