Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Scalable medication extraction and discontinuation identification from electronic health records using large language
Chong Shao1, Douglas Snyder2, Chiran Li2
1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA; Harvard T.H. Chan School of Public Health, Harvard University, Boston, MA, USA.
Objectives:
Identifying medication discontinuations in electronic health records (EHRs) is vital for patient safety but is often hindered by information being buried in unstructured notes. This study aims to evaluate the capabilities of advanced open-sourced and proprietary large language models in extracting medications and classifying their medication status from EHR notes, focusing on their scalability for medication information extraction without human annotation.
Study Design And Setting:
We collected three EHR datasets from diverse sources to build the evaluation benchmark: 1 publicly available dataset (Reannotated Clinical Acronym Sense Inventory dataset [Re-CASI]), 1 we annotated based on public MIMIC notes (MIMIC-IV Medication Snippet dataset [MIV-Med]), and 1 internally annotated on clinical notes from Mass General Brigham (MGB-Med). We evaluated 12 advanced LLMs, including general-domain open-sourced models (eg, Llama-3.1-70B-Instruct, Qwen2.5-72B-Instruct), medical-specific models (eg, MeLLaMA-70B-chat), and a proprietary model (GPT-4o). We explored multiple LLM prompting strategies, including zero-shot, 5-shot, and Chain-of-Thought (CoT) approaches. Performance on medication extraction, medication status classification, and their joint task (extraction then classification) was systematically compared across all experiments.
Results:
LLMs showed promising performance on medication extraction, while discontinuation classification and joint tasks were more challenging. GPT-4o consistently achieved the highest average F1 scores in all tasks under zero-shot setting - 94.0% for medication extraction, 78.1% for discontinuation classification, and 72.7% for the joint task. Open-sourced models followed closely, with Llama-3.1-70B-Instruct achieving the highest performance in medication status classification on the MIV-Med dataset (68.7%) and in the joint task on both the Re-CASI (76.2%) and MIV-Med (60.2%) datasets. Medical-specific LLMs demonstrated lower performance compared to advanced general-domain LLMs. Few-shot learning generally improved performance, while CoT reasoning showed inconsistent gains. Notably, open-sourced models occasionally surpassed GPT-4o performance, underscoring their potential in privacy-sensitive clinical research.
Conclusion:
LLMs demonstrate strong potential for medication extraction and discontinuation identification on EHR notes, with open-sourced models offering scalable alternatives to proprietary systems and few-shot learning further improving LLMs' capability.
Plain Language Summary:
Stopping a medicine can affect safety and treatment decisions, yet this detail is often buried in long electronic health record notes. We evaluated whether large language models, which read and summarize text, can automatically find medication names and decide whether each medicine is still being taken, has been stopped, or neither. We tested 12 models, including open-source options suitable for secure hospital use, on three collections of clinical notes and compared three simple instruction styles: giving no examples, showing a few examples, and asking for step-by-step reasoning. All models produced usable results. The strongest systems scored about 94 for finding medication names and about 78 for deciding continued or stopped status, on a standard 0 to 100 measure that balances completeness and correctness. Showing a few examples usually helped more than step-by-step prompts, and several open-source models performed close to a leading proprietary system. These tools could help hospitals and researchers monitor medications at scale to support drug-safety studies, adherence tracking, and clinical decision support, with local validation and safeguards before clinical use.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Clearance Models: Physiological Models
The organ's clearance rate depends on the blood flow to the organ and the extraction ratio (E). The extraction ratio describes the organ's...
Analysis of Population Pharmacokinetic Data