Related Experiment Video
Updated: Aug 27, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Omicron detection with large language models and YouTube audio data
James T Anibal1,2, Adam J Landa1, Nguyen T T Hang3
1Center for Interventional Oncology, Radiology and Imaging Sciences, NIH Clinical Center, National Cancer Institute, National Institute of Biomedical Imaging and Bioengineering, National Institutes of Health, 10 Center Dr, Building 10, Room 1C341, MSC 1182, Bethesda, MD 20892-1182 USA.
Large language models can detect COVID-19 from online audio data with high accuracy. This study shows the potential of using public audio for digital health and pandemic management tools.
Area of Science:
- Digital Health
- Artificial Intelligence
- Public Health
Background:
- Publicly available audio data offers a novel resource for developing digital health technologies.
- Large language models (LLMs) show promise in analyzing complex datasets for health applications.
Approach:
- Collected and transcribed YouTube audio data from individuals with self-declared COVID-19, upper respiratory infections (URI), and healthy controls.
- Utilized LLMs to detect self-reported COVID-19 cases, differentiate from other respiratory illnesses, and classify COVID-19 variants based on described symptoms.
- Employed the Whisper model for transcription and optimized prompts for LLM analysis.
Key Points:
- LLMs achieved 0.89 accuracy in identifying self-reported COVID-19 cases and 0.97 accuracy in distinguishing them from other respiratory illnesses.
- LLMs demonstrated a 0.77 mean accuracy in classifying COVID-19 variants using only symptom descriptions from audio data.
- This research contrasts with prior studies by using unscripted, real-world audio data from public online sources.
Conclusions:
- Publicly available audio data, analyzed by LLMs, presents a viable method for identifying COVID-19 and related respiratory conditions.
- This approach offers a new paradigm for pandemic management tools, highlighting the utility of audio data in clinical and public health surveillance.
- The findings underscore the potential of leveraging everyday online audio for scalable and accessible health monitoring.
More Related Videos
Related Concept Videos
Air-entraining Agents
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

