Related Experiment Video
Updated: Jun 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models outperform traditional natural language processing methods in extracting patient-reported
Perseus V Patel1,2, Conner Davis3, Amariel Ralbovsky1
1Department of Pediatrics, University of California San Francisco, San Francisco, CA.
Large language models (LLMs) accurately extract patient-reported outcomes (PROs) from inflammatory bowel disease (IBD) clinical notes. LLMs demonstrate superior generalizability and accuracy across institutions compared to traditional NLP methods.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Gastroenterology
Background:
- Patient-reported outcomes (PROs) are crucial for assessing inflammatory bowel disease (IBD) activity and treatment effectiveness.
- Manual extraction of PROs from clinical notes is time-consuming and labor-intensive.
- Improving data curation from electronic health records (EHRs) is essential for research and quality improvement.
Purpose of the Study:
- To compare traditional natural language processing (tNLP) with large language models (LLMs) for extracting IBD PROs.
- To evaluate the accuracy and generalizability of LLMs in extracting abdominal pain, diarrhea, and fecal blood from clinical notes.
- To assess the potential of LLMs in enhancing IBD research and patient care.
Main Methods:
- Annotation of clinic notes for three IBD PROs using predefined protocols.
- Development and internal testing of tNLP and LLM models at UCSF.
- External validation of models at Stanford University, comparing accuracy, sensitivity, specificity, and predictive values.
- Fairness and error assessments were conducted for all models.
Main Results:
- Inter-rater reliability for annotation exceeded 90%.
- On internal testing, LLMs (GPT-4) achieved accuracies of 96% (abdominal pain), 88% (diarrhea), and 90% (fecal blood), comparable to top tNLP models.
- External validation showed tNLP models failing to generalize (61-62% accuracy), while GPT-4 maintained >90% accuracy; PaLM-2 and GPT-4 performed similarly. No demographic biases were detected.
Conclusions:
- Large language models offer accurate and generalizable methods for extracting PROs from IBD clinical notes.
- LLMs maintain high accuracy across different institutions, irrespective of variations in note templates and authors.
- Widespread adoption of LLMs can significantly advance IBD research and improve patient care.
Related Concept Videos
Inflammatory Bowel Disease III: Diagnostic Studies and Management I-Nutritional Therapy
Diagnostic studies
A colonoscopy is the definitive screening test, distinguishing ulcerative colitis from other colon diseases with similar symptoms. During a colonoscopy test, inflamed mucosa with exudate ulcerations can be observed, and biopsies are taken to determine the histologic characteristics of the...
Inflammatory Bowel Disease IV: Pharmacological Management
Pharmacologic...
Drugs for Treatment of Crohn's Disease in IBD Using Glucocorticoids
Inflammatory Bowel Disease V: Surgical Management
Here are some common surgical interventions for IBD:
Inflammatory Bowel Disease II: Crohn's Disease
Inflammatory bowel disease, commonly known as IBD, refers to a collection of disorders that lead to persistent inflammation of the gastrointestinal tract. The two types of IBD are ulcerative colitis, which impacts the colon, and Crohn's disease, which can involve any part of the gastrointestinal segment.
Crohn's disease
Crohn's disease is a chronic, systemic inflammatory bowel disease (IBD) that predominantly affects the gastrointestinal tract. It is marked by...
Drugs for Treatment of Crohn's Disease in IBD Using Biologic Agents: Anti-TNF

