Related Experiment Video
Updated: Jun 5, 2025

03:53
Author Spotlight: Advancements in Multiplex Detection of Respiratory Viruses
Published on: November 10, 2023
1.1K
Discovering Time-Varying Public Interest for COVID-19 Case Prediction in South Korea Using Search Engine Queries:
Seong-Ho Ahn1, Kwangil Yim2, Hyun-Sik Won1
1Department of Artificial Intelligence, The Catholic University of Korea, Bucheon-Si, Republic of Korea.
Journal of Medical Internet Research
|December 16, 2024
Summary
This study introduces a new COVID-19 case prediction model that dynamically extracts relevant search queries using word embeddings. The model accurately forecasts cases by adapting to changing pandemic dynamics, outperforming previous methods.
Area of Science:
- Epidemiology
- Computational Linguistics
- Machine Learning
Background:
- Forecasting COVID-19 cases is vital for policy and lifestyle adjustments.
- Previous machine learning models for case prediction used static queries, limiting their ability to capture pandemic dynamics.
- Temporal variations in public search queries are crucial for understanding and predicting pandemic trends.
Purpose of the Study:
- To develop a novel framework for extracting COVID-19-related keywords that account for temporal variations.
- To investigate time-delayed web search behavior and its correlation with public interest in COVID-19.
- To improve COVID-19 case prediction by incorporating dynamically extracted, time-sensitive keywords.
Main Methods:
- Trained word embedding models on a news corpus to extract time-varying keywords related to "Corona" over 4-month intervals.
- Utilized time-lagged cross-correlation to identify optimal time lags between expanded queries and confirmed COVID-19 cases.
- Applied Principal Component Analysis (PCA) for feature reduction and ElasticNet regression for predicting daily case counts.
Main Results:
- Successfully extracted phase-specific keywords reflecting COVID-19 symptoms, societal impact, policy responses, and public sentiment (e.g., "economic crisis", "anxiety").
- Models trained with time-lagged, dynamically extracted queries significantly outperformed previous methods for 1-14 day ahead predictions.
- Demonstrated superior performance compared to models using only past case counts or static queries, particularly for 9-11 day ahead predictions (P<.01).
Conclusions:
- A novel COVID-19 case prediction model was developed using automated, time-aware keyword extraction via word embedding.
- The proposed model surpasses traditional methods relying on static or heuristic queries, requiring no prior expert knowledge.
- The approach effectively captures temporal shifts in public interest, offering a dynamic tool for pandemic monitoring and prediction.
Keywords:
COVID-19South Koreacase predictionconfirmed case predictioninfodemiologyinfodemiology studylifestylemachine learningmachine learning techniquesmodelnovel frameworkpolicyprediction modelpublic healthquery expansionsearch enginesearch engine queriestemporaltemporal semanticstemporal variationutilizationweb-based searchword embeddingMore Related Videos
Related Concept Videos
Steps in Outbreak Investigation
105
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
105
Factorial Design
13.0K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.0K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K

