发现韩国使用搜索引擎查询对COVID-19病例预测的时间变化的公众兴趣:传染病学研究
Seong-Ho Ahn1, Kwangil Yim2, Hyun-Sik Won1
1Department of Artificial Intelligence, The Catholic University of Korea, Bucheon-Si, Republic of Korea.
Journal of medical Internet research
|December 16, 2024
概括
这项研究引入了一种新的COVID-19病例预测模型,该模型使用词嵌入动态提取相关的搜索查询. 该模型通过适应不断变化的流行病动态,准确地预测病例,优于以前的方法.
科学领域:
- 流行病学 流行病学
- 计算语言学 计算语言学
- 机器学习 机器学习
背景情况:
- 预测COVID-19病例对于政策和生活方式调整至关重要.
- 之前用于病例预测的机器学习模型使用了静态查询,限制了它们捕捉大流行动态的能力.
- 公共搜索查询的时间变化对于理解和预测流行病趋势至关重要.
研究的目的:
- 开发一个新的框架来提取与COVID-19相关的关键词,以解释时间变化.
- 调查时间延迟的网络搜索行为及其与 COVID-19 公众兴趣的相关性.
- 通过结合动态提取的,时间敏感的关键字来改善COVID-19病例预测.
主要方法:
- 在新闻群体上训练有素的词嵌入模型,在4个月的时间间隔内提取与"冠状病毒"相关的时间变化的关键词.
- 利用时间延迟交叉相关性来识别扩展查询和确诊的COVID-19病例之间的最佳时间延迟.
- 应用主要组件分析 (PCA) 用于特征减少和ElasticNet回归来预测每日病例数.
主要成果:
- 成功提取了反映COVID-19症状,社会影响,政策反应和公众情绪 (例如",经济危机"",焦虑") 的特定阶段关键词.
- 用时间滞后,动态提取查询训练的模型在1-14天前的预测中显著超过了以前的方法.
- 与仅使用过去病例计数或静态查询的模型相比,表现出优异的性能,特别是对于9-11天前的预测 (P<.01).
结论:
- 开发了一种新的COVID-19病例预测模型,使用通过词嵌入实现自动化,时间意识的关键词提取.
- 拟议的模型超越了依赖于静态或启发式查询的传统方法,不需要事先的专家知识.
- 该方法有效地捕捉了公共利益的时间变化,为流行病监测和预测提供了一个动态的工具.
关键词:
在 COVID-19 疫情中,韩国 韩国 韩国 韩国案例预测情况预测.确诊病例预测 确诊病例预测传染病学是传染病学.信息传染病学研究研究生活方式 生活方式机器学习是机器学习.机器学习技术 机器学习技术模型模型模型模型模型模型这是一个新的框架.政策 政策 政策 政策预测模型 预测模型公共卫生公共卫生.查询扩展 查询扩展搜索引擎搜索引擎搜索引擎是什么搜索引擎查询中的查询.时间 时间 时间 时间时间语义的时间语义.时间变化的时间变化.使用利用利用利用利用利用利用利用利用利用基于网络的搜索搜索.一个词嵌入的词嵌入.更多相关视频
相关概念视频
Steps in Outbreak Investigation
105
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
105
Factorial Design
13.0K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.0K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K


