Related Experiment Video
Updated: May 15, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Integrating large language models with human expertise for disease detection in electronic health records
Jie Pan1, Seungwon Lee2, Cheligeer Cheligeer3
1Centre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary, AB, Canada; Department of Community Health Sciences, Cumming School of Medicine, University of Calgary, Calgary, AB, Canada; Libin Cardiovascular Institute, University of Calgary, Calgary, AB, Canada.
Objective:
Electronic health records (EHR) are widely available to complement administrative data-based disease surveillance and healthcare performance evaluation. Defining conditions from EHR is labour-intensive and requires extensive manual labelling of disease outcomes. This study developed an efficient strategy based on advanced large language models to identify multiple conditions from EHR clinical notes.
Methods:
We linked a cardiac registry cohort in 2015 with an EHR system in Alberta, Canada. We developed a pipeline that leveraged a generative large language model (LLM) to analyze, understand, and interpret EHR notes by prompts based on specific diagnosis, treatment management, and clinical guidelines. The pipeline was applied to detect acute myocardial infarction (AMI), diabetes, and hypertension. The performance was compared against clinician-validated diagnoses as the reference standard and widely adopted International Classification of Diseases (ICD) codes-based methods.
Results:
The study cohort accounted for 3088 patients and 551,095 clinical notes. The prevalence was 55.4 %, 27.7 %, 65.9 % and for AMI, diabetes, and hypertension, respectively. The performance of the LLM-based pipeline for detecting conditions varied: AMI had 88 % sensitivity, 63 % specificity, and 77 % positive predictive value (PPV); diabetes had 91 % sensitivity, 86 % specificity, and 71 % PPV; and hypertension had 94 % sensitivity, 32 % specificity, and 72 % PPV. Compared with ICD codes, the LLM-based method demonstrated improved sensitivity and negative predictive value across all conditions. The monthly percentage trends from the detected cases by LLM and reference standard showed consistent patterns.
Conclusion:
The proposed LLM-based pipeline demonstrated reasonable accuracy and high efficiency in disease detection for multiple conditions. Human expert knowledge can be integrated into the pipeline to guide EHR note analysis without manually curated labels. The method could enable comprehensive real-time disease surveillance using EHRs.
Related Concept Videos
Methods of Documentation VII: EMR
Steps in Outbreak Investigation
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...

