Related Experiment Video
Updated: Jun 7, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging Open-Source Large Language Models for Data Augmentation in Hospital Staff Surveys: Mixed Methods Study
Carl Ehrett1, Sudeep Hegde2, Kwame Andre3
1Watt Family Innovation Center, Clemson University, Clemson, SC, United States.
Open-source large language models (LLMs) can effectively augment small healthcare datasets for text classification. This approach enhances classifier performance, offering privacy-conscious solutions for medical education and patient care.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing
- Medical Education Technology
Background:
- Generative large language models (LLMs) show potential for medical education but their use in healthcare for augmenting small datasets, especially with privacy and cost constraints, is underexplored.
- Existing LLM applications often rely on third-party services, limiting their use in sensitive healthcare contexts.
Purpose of the Study:
- To investigate the efficacy of open-source LLMs for data augmentation in text classification tasks within healthcare.
- To evaluate the performance of models like Large Language Model Meta AI (LLaMA) and Alpaca for generating synthetic data for hospital staff surveys.
Main Methods:
- A two-step process involving data augmentation and text classification was employed.
- Four open-source generative LLMs were used to create synthetic data from hospital staff surveys concerning COVID-19 pandemic adaptations.
- Three distinct classifier LLMs were then used to categorize the augmented text data.
Main Results:
- The best performance was achieved using LLaMA 7B (temperature 0.7, 100 augments) for data augmentation and Robustly Optimized BERT Pretraining Approach (RoBERTa) for classification, yielding an average AUC of 0.87.
- Open-source LLMs significantly improved text classifier performance on limited healthcare datasets.
Conclusions:
- Open-source LLMs offer a viable solution for data augmentation in healthcare settings, enhancing text classification accuracy.
- The study underscores the importance of privacy and ethical considerations when implementing LLMs in medical applications.
- Future research should explore further applications and optimizations of LLMs in medical education and patient care.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Surveys
Data Collection by Survey
Statistical Software for Data Analysis and Clinical Trials
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Statistical Methods for Analyzing Epidemiological Data

