Related Experiment Video
Updated: Jul 30, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Model as Unsupervised Health Information Retriever.
Keyuan Jiang1, Mohammed M Mujtaba1, Gordon R Bernard2
1Purdue University Northwest, Hammond, Indiana, USA.
This study shows that large language models can find COVID-19 symptoms in social media posts without prior examples. This zero-shot learning approach efficiently identifies health information and aids future research.
Area of Science:
- Natural Language Processing
- Public Health Informatics
- Computational Linguistics
Background:
- Accessing reliable health information is crucial for disease management.
- Self-reported health data, particularly from social media, can offer valuable insights into disease symptoms and trends.
- COVID-19 has highlighted the need for efficient methods to track public health information.
Purpose of the Study:
- To evaluate the effectiveness of a pretrained large language model (GPT-3) in retrieving COVID-19 symptom mentions from Twitter data.
- To assess the utility of a zero-shot learning approach for identifying health-related information without manual data annotation.
- To introduce and utilize a novel performance metric, total match (TM), encompassing exact, partial, and semantic matches.
Main Methods:
- Utilized a pretrained large language model (GPT-3) for zero-shot learning to identify symptom mentions in COVID-19-related Twitter posts.
- Developed and applied a 'total match' (TM) metric to evaluate the accuracy of symptom identification, considering various match types.
- Compared the performance of the zero-shot method against the need for annotated datasets.
Main Results:
- The zero-shot learning approach demonstrated significant capability in retrieving symptom mentions from Twitter data.
- The 'total match' metric provided a comprehensive evaluation of the model's performance in identifying health information.
- The study confirmed that zero-shot learning is a powerful technique that does not require data annotation.
Conclusions:
- Zero-shot learning with large language models is an effective strategy for extracting health information, specifically COVID-19 symptoms, from social media.
- This method reduces the dependency on manually annotated datasets, accelerating health information retrieval.
- The findings suggest that zero-shot learning can serve as a foundational step for generating data for few-shot learning, potentially enhancing performance further.
More Related Videos
Related Concept Videos
Health Literacy
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Models of Health Promotion and Illness Prevention II
The agent-host-environment model states that disease results...

