预测加利福尼亚州的COVID-19新病例,使用谷歌趋势数据和机器学习方法
Amir Habibdoust1, Maryam Seifaddini2, Moosa Tatar3
1Institute for Data Science and Informatics, University of Missouri, Columbia, Missouri, USA.
Informatics for health & social care
|February 14, 2024
概括
谷歌趋势数据通过分析"发烧"和"COVID测试"等搜索术语,准确地预测COVID-19病例. 这种方法补充了现有的流行病学模型,以获得更好的公共卫生见解.
科学领域:
- 流行病学 流行病学
- 数据科学数据科学数据科学
- 公共卫生 公共卫生
背景情况:
- 谷歌趋势为预测传染病趋势提供了有价值的数据.
- 实时数据源对于监测公共卫生危机至关重要.
研究的目的:
- 评估谷歌趋势数据在预测加利福尼亚州COVID-19新病例中的准确性.
- 使用谷歌趋势数据,将GMDH型神经网络模型与LSTM模型的预测性能进行比较.
主要方法:
- 开发并使用了GMDH类型的神经网络模型.
- 分析了三个不同时期的谷歌查询数据 (2020年3月至7月,2020年10月至2021年1月,2020年5月至2021年1月).
- 将模型性能与长期短期记忆 (LSTM) 模型进行比较.
主要成果:
- 谷歌相对搜索量 (RSV) 准确地预测了新的COVID-19病例.
- 特定的搜索术语,如"发烧"",COVID测试"",COVID症状"",COVID治疗"",呼吸短促"显著提高了预测准确度.
结论:
- 谷歌趋势数据为检测COVID-19趋势提供了一个有价值的,近乎实时的资源.
- 这些数据可以补充和扩展传统的流行病学模型.
- 利用搜索查询数据提高了追踪和预测传染病爆发的能力.
更多相关视频
10:46A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
10.7K
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
7.1K
相关概念视频
Steps in Outbreak Investigation
128
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
128
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Statistical Methods for Analyzing Epidemiological Data
366
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
366
End Point Prediction: Gran Plot
324
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
324
Aggregates Classification
325
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
325
