Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Extracting TNFi switching reasons and trajectories from real-world data using large language models
Brenda Y Miao1, Marie Binvignat1,2,3, Augusto Garcia-Agundez4
1Bakar Computational Health Sciences Institute, University of California San Francisco, San Francisco, CA 94143, United States.
Background And Significance:
To evaluate whether large language models (LLMs) can automate chart review to identify tumor necrosis factor inhibitor (TNFi) switching patterns and reasons for switching in a large real-world cohort.
Materials And Methods:
We conducted an observational study using de-identified electronic health record (EHR) data from 2012 to 2023 at a single academic medical center (University of California, San Francisco). TNFi medication orders and linked clinical notes were extracted, requiring at least 6 months of follow-up to identify treatment switches, defined as a change from one TNFi to another at consecutive encounters. Using GPT-4, we extracted which TNFi was stopped and started and classified the reason for switching. Performance was benchmarked against eight open-source LLMs, structured EHR data, and expert annotation.
Results:
A total of 9187 patients (mean [SD] age, 39.9 [19.0] years; 57.1% female) received ≥1 TNFi with sufficient follow-up. We identified 3104 TNFi switches among 2112 patients. GPT-4 achieved micro-F1 scores of 0.75 (stopped drug), 0.80 (started drug), and 0.83 (reason). Among open-source models, Starling-7B-beta and Llama-3-8B performed most competitively. The most common reason identified by GPT-4 was lack of efficacy (56.9%), followed by adverse events (13.5%) and insurance/cost (10.8%).
Conclusions:
Both GPT-4 and locally deployable LLMs effectively extracted complex treatment trajectories and rationale from clinical notes, supporting their broader utility in scalable EHR review and real-world evidence generation.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language and Cognition
Leaky Scanning