Predicting COVID-19 new cases in California with Google Trends data and a machine learning approach
Amir Habibdoust1, Maryam Seifaddini2, Moosa Tatar3
1Institute for Data Science and Informatics, University of Missouri, Columbia, Missouri, USA.
Background:
Google Trends data can be a valuable source of information for health-related issues such as predicting infectious disease trends.
Objectives:
To evaluate the accuracy of predicting COVID-19 new cases in California using Google Trends data, we develop and use a GMDH-type neural network model and compare its performance with a LTSM model.
Methods:
We predicted COVID-19 new cases using Google query data over three periods. Our first period covered March 1, 2020, to July 31, 2020, including the first peak of infection. We also estimated a model from October 1, 2020, to January 7, 2021, including the second wave of COVID-19 and avoiding possible biases from public interest in searching about the new pandemic. In addition, we extended our forecasting period from May 20, 2020, to January 31, 2021, to cover an extended period of time.
Results:
Our findings show that Google relative search volume (RSV) can be used to accurately predict new COVID-19 cases. We find that among our Google relative search volume terms, "Fever," "COVID Testing," "Signs of COVID," "COVID Treatment," and "Shortness of Breath" increase model predictive accuracy.
Conclusions:
Our findings highlight the value of using data sources providing near real-time data, e.g., Google Trends, to detect trends in COVID-19 cases, in order to supplement and extend existing epidemiological models.
More Related Videos
10:46A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Steps in Outbreak Investigation
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Methods for Analyzing Epidemiological Data
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
