Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Steps in Outbreak Investigation01:18

Steps in Outbreak Investigation

350
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
350
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

723
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
723
Principles of Disease Surveillance01:26

Principles of Disease Surveillance

329
Disease surveillance is the systematic collection, analysis, and interpretation of health data essential to the planning, implementation, and evaluation of public health practice. This process integrates data dissemination to entities responsible for preventing and controlling disease, injury, and disability. Surveillance systems provide crucial information for action, helping public health authorities make informed decisions to manage and prevent outbreaks, ensure public safety, optimize...
329
Data Collection by Observations01:08

Data Collection by Observations

14.1K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
14.1K
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

1.2K
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Reddit posts reveal how natural environments affect social anxiety in young people.

medRxiv : the preprint server for health sciences·2026
Same author

Advanced topic modeling with large language models: analyzing social media content from dementia caregivers.

Innovation in aging·2025
Same author

WATCH-SS: Developing a Trustworthy and Explainable Modular Framework for Detecting Cognitive Impairment from Spontaneous Speech.

medRxiv : the preprint server for health sciences·2025
Same author

Automatic genetic phenotype normalization from dysmorphology physical examinations: an overview of the BioCreative VIII-Task 3 competition.

Database : the journal of biological databases and curation·2025
Same author

Health-Related Concerns of Anti-LGBTQ+ Legislation: Thematic Analysis Using Social Media Data.

JMIR infodemiology·2025
Same author

Association Between COVID-19 During Pregnancy and Preterm Birth by Trimester of Infection: Retrospective Cohort Study Using Large-Scale Social Media Data.

Journal of medical Internet research·2025

Related Experiment Video

Updated: Nov 21, 2025

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs
07:13

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs

Published on: April 9, 2021

4.4K

Toward Using Twitter for Tracking COVID-19: A Natural Language Processing Pipeline and Exploratory Data Set.

Ari Z Klein1, Arjun Magge1, Karen O'Connor1

  • 1Department of Biostatistics, Epidemiology, and Informatics, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, United States.

Journal of Medical Internet Research
|January 15, 2021
PubMed
Summary

Researchers developed an automated system using Twitter data to identify potential COVID-19 cases in the US, complementing traditional testing methods. This approach helps track the pandemic

Keywords:
COVID-19coronavirusdata miningepidemiologyinfodemiologynatural language processingpandemicssocial media

More Related Videos

Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples
09:26

Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples

Published on: June 30, 2023

1.4K
Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
06:22

Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections

Published on: September 19, 2025

198

Related Experiment Videos

Last Updated: Nov 21, 2025

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs
07:13

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs

Published on: April 9, 2021

4.4K
Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples
09:26

Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples

Published on: June 30, 2023

1.4K
Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
06:22

Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections

Published on: September 19, 2025

198

Area of Science:

  • Computational epidemiology
  • Public health surveillance
  • Natural Language Processing (NLP)

Background:

  • COVID-19 spread monitoring is challenged by testing limitations in the US.
  • Testing shortages and delays hinder accurate real-time case tracking.
  • Need for complementary surveillance methods beyond traditional testing.

Purpose of the Study:

  • Develop and deploy an automated NLP pipeline to analyze Twitter data.
  • Identify potential COVID-19 cases not captured by official testing.
  • Create a complementary resource for tracking disease spread.

Main Methods:

  • Collected and processed English tweets mentioning COVID-19 keywords via Twitter API.
  • Used regular expressions and NLP to filter tweets for self-reported potential cases.
  • Annotated tweets and trained a Bidirectional Encoder Representations from Transformers (BERT) deep neural network classifier.
  • Deployed the pipeline on over 85 million tweets for case identification.

Main Results:

  • Achieved an F1-score of 0.76 in detecting self-reported COVID-19 cases using the BERT model.
  • Identified 13,714 tweets self-reporting potential COVID-19 cases with US state-level geolocations.
  • High interannotator agreement (Cohen κ=0.77) validated annotation quality.

Conclusions:

  • Publicly released 13,714 geolocated tweets indicating potential COVID-19 cases.
  • Twitter data offers a valuable complementary resource for tracking COVID-19.
  • Facilitates future research on social media's role in public health surveillance.