Related Experiment Video
Updated: Jun 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Generative AI and unstructured audio data for precision public health
James Anibal1,2, Adam Landa1, Hang Nguyen3
1Center for Interventional Oncology, Radiology and Imaging Sciences, NIH Clinical Center, Bethesda, USA.
Abstract:
In this study, transcribed videos about personal experiences with COVID-19 were used for variant classification. The o1 LLM was used to summarize the transcripts, excluding references to dates, vaccinations, testing methods, and other variables that were correlated with specific variants but unrelated to changes in the disease. This step was necessary to effectively simulate model deployment in the early days of a pandemic when subtle changes in symptomatology may be the only viable biomarkers of disease mutations. The embedded summaries were used for training a neural network to predict the variant status of the speaker as "Omicron" or "Pre-Omicron", resulting in an AUROC score of 0.823. This was compared to a neural network model trained on binary symptom data, which obtained a lower AUROC score of 0.769. Results of the study illustrated the future value of LLMs and audio data in the design of pandemic management tools for health systems.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Non-equilibrium in the Cell
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Steps in Outbreak Investigation
Sampling Methods: Overview
In analytical chemistry, the choice of...
Genome Annotation and Assembly

