Related Experiment Video
Updated: Apr 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the quality and performance of synthetic data augmentation to identify stigmatizing language in obstetric
Jihye Kim Scroggins1, Veronica Barcelona2, Ismael I Hulchafo2
1School of Nursing, University of North Carolina at Chapel Hill, Chapel Hill, NC.
Background:
Stigmatizing language in clinical notes is associated with adverse birth outcomes. Identifying stigmatizing language using natural language processing (NLP) requires human-annotated datasets, which can be scarce.
Purpose:
Generate synthetic data containing stigmatizing language using Llama-3, assess quality, and examine NLP performance.
Methods:
Synthetic data was generated for "Difficult Patient" and "Questioning Patient Credibility" using zero-shot (no examples), few-shot (examples), and fine-tuned few-shot (fine-tuned model and examples). Nurse researchers evaluated data quality. NLP models (e.g., ClinicalBERT) were trained using real-world (birth admission notes) and synthetic augmented datasets.
Findings:
Fine-tuned few-shot generated higher quality synthetic data. ClinicalBERT F1-scores increased from 0.88 to 0.95 with 10% synthetic data augmentation for Difficult Patient and from 0.68 to 0.82 with 50% augmentation for Questioning Patient Credibility.
Discussion:
Synthetic data improved performance with varied gains, showing potential to address data scarcity. Findings support developing tools to identify stigmatizing language, providing measurable indicators to inform policy.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Genetic Lingo

