Related Experiment Video
Updated: Jul 10, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
From Zero-Shot to Bedside: A Practical Playbook for Adapting Open-Source Large Language Models to Clinical Symptom
Li-Ching Chen1,2,3, Travis Zack2,3,4, Divneet Mandair2
1UC Berkeley.
Summary
This study offers a practical guide for fine-tuning large language models (LLMs) on clinical notes. An LLM-assisted workflow improved annotation accuracy and reduced expert review burden for pancreatic cancer patients.
Area of Science:
- Clinical Natural Language Processing (NLP)
- Artificial Intelligence in Medicine
- Machine Learning for Healthcare
Background:
- Large language models (LLMs) show promise for analyzing clinical notes, but practical guidance on adapting open-source models and ensuring annotation quality is scarce.
- Fine-tuning LLMs on sensitive clinical data requires careful consideration of privacy and task-specific adaptation.
Purpose of the Study:
- To provide a playbook for fine-tuning open-source LLMs on de-identified clinical notes for pancreatic cancer patients.
- To evaluate different prompting strategies and compare open-source models with proprietary ones like GPT-4o.
- To develop and assess an LLM-assisted adjudication workflow for improving annotation quality and reducing expert burden.
Main Methods:
- Fine-tuning of open-source LLMs on de-identified clinical notes from pancreatic cancer patients (pre-diagnosis and on-treatment).
- Evaluation of various prompting strategies and disease-level vs. task-specific adaptation.
- Implementation of an LLM-assisted adjudication workflow to flag conflicting predictions for expert review.
- Assessment of machine-generated annotations to augment limited expert labels.
Main Results:
- The LLM-assisted adjudication workflow effectively identified annotation errors, concentrating expert review on a small subset of notes and improving downstream model performance.
- Using a balanced mix of synthetic and human data for training enhanced the performance of fine-tuned models.
- Strategies were identified to improve accuracy and reduce the annotation burden when deploying LLMs in clinical settings.
Conclusions:
- This work offers practical strategies for adapting and deploying open-source LLMs in clinical NLP tasks, enhancing accuracy and efficiency.
- The developed LLM-assisted adjudication workflow is a key innovation for managing annotation quality at scale.
- The findings support privacy-preserving, site-adapted clinical NLP through effective LLM fine-tuning and data augmentation.