Related Experiment Video
Updated: May 30, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models Outperform Traditional Natural Language Processing Methods in Extracting Patient-Reported
Perseus V Patel1,2, Conner Davis3, Amariel Ralbovsky1
1Department of Pediatrics, University of California, San Francisco, San Francisco, California.
Background And Aims:
Patient-reported outcomes (PROs) are vital in assessing disease activity and treatment outcomes in inflammatory bowel disease (IBD). However, manual extraction of these PROs from the free-text of clinical notes is burdensome. We aimed to improve data curation from free-text information in the electronic health record, making it more available for research and quality improvement. This study aimed to compare traditional natural language processing (tNLP) and large language models (LLMs) in extracting 3 IBD PROs (abdominal pain, diarrhea, fecal blood) from clinical notes across 2 institutions.
Methods:
Clinic notes were annotated for each PRO using preset protocols. Models were developed and internally tested at the University of California, San Francisco, and then externally validated at Stanford University. We compared tNLP and LLM-based models on accuracy, sensitivity, specificity, positive, and negative predictive value. In addition, we conducted fairness and error assessments.
Results:
Interrater reliability between annotators was >90%. On the University of California, San Francisco test set (n = 50), the top-performing tNLP models showcased accuracies of 92% (abdominal pain), 82% (diarrhea) and 80% (fecal blood), comparable to GPT-4, which was 96%, 88%, and 90% accurate, respectively. On external validation at Stanford (n = 250), tNLP models failed to generalize (61%-62% accuracy) while GPT-4 maintained accuracies >90%. Pathways Language Model-2 and Generative Pre-trained Transformer-4 showed similar performance. No biases were detected based on demographics or diagnosis.
Conclusion:
LLMs are accurate and generalizable methods for extracting PROs. They maintain excellent accuracy across institutions, despite heterogeneity in note templates and authors. Widespread adoption of such tools has the potential to enhance IBD research and patient care.
More Related Videos
09:44Functional Assessment of Intestinal Motility and Gut Wall Inflammation in Rodents: Analyses in a Standardized Model of Intestinal Manipulation
Published on: September 11, 2012
09:04DNBS/TNBS Colitis Models: Providing Insights Into Inflammatory Bowel Disease and Effects of Dietary Fat
Published on: February 27, 2014
Related Concept Videos
Inflammatory Bowel Disease III: Diagnostic Studies and Management I-Nutritional Therapy
Diagnostic studies
A colonoscopy is the definitive screening test, distinguishing ulcerative colitis from other colon diseases with similar symptoms. During a colonoscopy test, inflamed mucosa with exudate ulcerations can be observed, and biopsies are taken to determine the histologic characteristics of the...
Inflammatory Bowel Disease IV: Pharmacological Management
Pharmacologic...
Drugs for Treatment of Crohn's Disease in IBD Using Glucocorticoids
Inflammatory Bowel Disease V: Surgical Management
Here are some common surgical interventions for IBD:
Inflammatory Bowel Disease II: Crohn's Disease
Inflammatory bowel disease, commonly known as IBD, refers to a collection of disorders that lead to persistent inflammation of the gastrointestinal tract. The two types of IBD are ulcerative colitis, which impacts the colon, and Crohn's disease, which can involve any part of the gastrointestinal segment.
Crohn's disease
Crohn's disease is a chronic, systemic inflammatory bowel disease (IBD) that predominantly affects the gastrointestinal tract. It is marked by...
Drugs for Treatment of Crohn's Disease in IBD Using Biologic Agents: Anti-TNF