Related Experiment Video
Updated: Jul 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Linguistic Fidelity and Classification Performance of Large Language Models for Generating Synthetic Operative Notes:
Meredith Cox1, Elaine Lin1, Nicholas Oleck1
1Division of Plastic, Oral, and Maxillofacial Surgery, Duke University Hospital, 2301 Erwin Road, Durham, NC, 27710, United States, 1 919-668-3110.
Large language models (LLMs) can generate synthetic surgical notes with high linguistic fidelity. This synthetic data improves machine learning model performance, especially when real surgical data is scarce.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Machine Learning in Surgery
Background:
- Machine learning for surgical applications requires extensive, diverse datasets, but data scarcity is a significant hurdle due to privacy, institutional variations, and rare procedures.
- Large language models (LLMs) present a potential solution for synthetic data generation, yet their effectiveness in specialized surgical fields needs further investigation.
Purpose of the Study:
- To evaluate the linguistic fidelity of LLM-generated operative notes for cleft lip and palate surgeries.
- To assess the impact of synthetic data on natural language processing (NLP) classifier performance under limited data conditions.
Main Methods:
- Generated synthetic operative notes using GPT-4o from 630 authentic cleft procedure records (lip repair, palate repair, alveolar bone grafting).
- Assessed linguistic fidelity using BERTScore (semantic similarity), Jensen-Shannon divergence (syntactic structure), and BLEU scores (lexical overlap).
- Trained NLP classifiers on real data under full and data-scarce (5-10% positive cases) conditions, with and without synthetic data augmentation.
Main Results:
- Synthetic notes showed high semantic fidelity (BERTScore F1: 0.86-0.88) and low syntactic divergence (JSD: 0.06-0.08).
- NLP classifier performance improved with synthetic data augmentation under data-scarce conditions (e.g., cleft lip AUC increased from 0.915 to 0.929 at 5% data).
- Augmentation had minimal impact on classifier performance with larger datasets (10% positive cases) or full data availability.
Conclusions:
- LLM-generated operative notes possess strong semantic and syntactic fidelity, comparable to authentic records.
- Synthetic data effectively enhances machine learning model performance in data-limited surgical scenarios, particularly for rare procedures.
- LLM-based synthetic data generation offers a viable strategy to overcome data scarcity in specialized surgical domains, facilitating robust ML model development.
Related Concept Videos
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Improving Translational Accuracy
Improving Translational Accuracy