Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Steps in Outbreak Investigation01:18

Steps in Outbreak Investigation

171
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
171
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

7.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.8K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

493
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
493
Pareto Chart00:52

Pareto Chart

6.9K
A Pareto chart is a bar graph or a combination of both line and bar graphs. The bar lengths represent the individual values or the frequency, while the lines represent the cumulative total values. In this chart, the longest bars are arranged on the left and the shortest bars on the right, which makes it easier to read and interpret the data. It can also be called a Pareto diagram or Pareto analysis.
The Pareto chart is named after the Italian economist Vilfredo Pareto, who described the Pareto...
6.9K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

15.6K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.6K
Issues And Trends In Healthcare Delivery System01:29

Issues And Trends In Healthcare Delivery System

5.8K
The issues and trends in healthcare delivery are constantly changing. The COVID-19 pandemic is one recent issue that wreaked havoc on healthcare systems, causing a shortage of healthcare workers, high demand for medicines and supplies, and increased medical expenditure due to a lack of insurance. Other issues include rising healthcare costs and care fragmentation.
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
5.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Predicting early and complete drug release from long-acting injectables using explainable machine learning.

International journal of pharmaceutics·2026
Same author

Predicting Early and Complete Drug Release from Long-Acting Injectables Using Explainable Machine Learning.

ArXiv·2026
Same author

Attention-based Imputation of Missing Values in Electronic Health Records Tabular Data.

Proceedings. IEEE International Conference on Healthcare Informatics·2024
Same author

Deep Clustering of Electronic Health Records Tabular Data for Clinical Interpretation.

... IEEE International Conference on Telecommunications and Photonics. IEEE International Conference on Telecommunications and Photonics·2024
Same author

Deep imputation of missing values in time series health data: A review with benchmarking.

Journal of biomedical informatics·2023
Same author

Perturbation of deep autoencoder weights for model compression and classification of tabular data.

Neural networks : the official journal of the International Neural Network Society·2022

Related Experiment Video

Updated: Aug 27, 2025

Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
08:53

Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community

Published on: May 31, 2019

5.2K

Mining Social Media Data to Predict COVID-19 Case Counts.

Maksims Kazijevs1, Furkan A Akyelken2, Manar D Samad2

  • 1Dept. of Computer Science, Tennessee State University, Nashville, TN, USA.

Proceedings. IEEE International Conference on Healthcare Informatics
|September 23, 2022
PubMed
Summary

Predicting COVID-19 case counts is possible using natural language processing (NLP) to analyze tweets. This method offers a novel way to forecast disease spread and inform public health interventions.

Keywords:
LSTMTwitternatural language processingpandemic predictionsocial media

More Related Videos

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs
07:13

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs

Published on: April 9, 2021

4.3K
Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples
09:26

Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples

Published on: June 30, 2023

1.2K

Related Experiment Videos

Last Updated: Aug 27, 2025

Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
08:53

Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community

Published on: May 31, 2019

5.2K
Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs
07:13

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs

Published on: April 9, 2021

4.3K
Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples
09:26

Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples

Published on: June 30, 2023

1.2K

Area of Science:

  • Computational epidemiology
  • Natural Language Processing (NLP)
  • Machine Learning

Background:

  • The COVID-19 pandemic's unpredictability necessitates advanced forecasting methods.
  • Existing prediction models often rely on traditional epidemiological data.
  • Social media data offers a rich, real-time source for public health insights.

Purpose of the Study:

  • To develop and evaluate an NLP-based model for predicting COVID-19 case counts (CCC).
  • To assess the efficacy of encoding tweets for time-series forecasting of disease spread.
  • To explore the potential of social media text for early pandemic detection.

Main Methods:

  • Utilized state-of-the-art NLP algorithms to numerically encode COVID-19 related tweets from eight US cities.
  • Proposed a city-embedding technique for time-series representation of daily tweets.
  • Employed a custom long-short term memory (LSTM) model for CCC prediction.

Main Results:

  • The universal sentence encoder achieved the best performance with a normalized root mean squared error (NRMSE) of 0.090 for predicting CCC six days ahead.
  • Achieved R-squared (R²) scores consistently above 0.70, often exceeding 0.8, indicating strong correlation between predicted and actual CCC.
  • Demonstrated robust performance across different cities and time-series data lengths.

Conclusions:

  • NLP-encoded tweet semantics can be effectively mapped to predict future COVID-19 case counts.
  • Social media text mining provides a valuable, direct method for tracking pandemic trajectories.
  • This approach can aid in early preventive measures and public health response planning.