Related Experiment Video
Updated: Jun 5, 2025

Analyzing Long-Term Electrocardiography Recordings to Detect Arrhythmias in Mice
Published on: May 23, 2021
Engineering of Generative Artificial Intelligence and Natural Language Processing Models to Accurately Identify
Ruibin Feng1, Kelly A Brennan1, Zahra Azizi1
1Department of Medicine, Stanford University, CA (R.F., K.A.B., Z.A., J.G., B.D., H.J.C., P.G., P.C., M. Pedron, S.R.-C., Y.B.D., H.D.L., T.B., M. V. P, M.R., A.J.R., S.M.N.).
Prompt engineering significantly improves large language models' (LLMs) accuracy in interpreting electronic health records (EHRs) for diagnostics. This technique enhances LLM performance without requiring specialized medical knowledge, making EHR data more accessible.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing for Healthcare
Background:
- Large language models (LLMs) excel with public data but struggle with private sources like electronic health records (EHRs).
- Prompt engineering is a potential method to enhance LLM accuracy for EHR data interpretation without domain expertise.
Purpose of the Study:
- To systematically test prompt engineering techniques for improving LLM accuracy in interpreting EHRs for nuanced diagnostic questions.
- To evaluate if prompt engineering can enable LLMs to identify clinical endpoints from EHRs with expert-level accuracy.
Main Methods:
- Designed and tested prompt engineering strategies on GPT-4-turbo using 490 EHR notes from 125 patients with heart rhythm disorders.
- Compared LLM performance against rule-based NLP and BERT-based models for identifying recurrent arrhythmias.
- Validated findings across GPT-3.5-turbo and Jurassic-2 LLMs.
Main Results:
- Out-of-the-box GPT-4-turbo accuracy was 64.3%, increasing to 91.4% with prompt engineering (rationale, structured output, exemplars).
- Prompt-engineered LLM performance significantly surpassed traditional NLP and BERT-based models (P<0.05).
- Consistent accuracy improvements were observed across different LLMs.
Conclusions:
- Prompt engineering enables LLMs to accurately identify clinical endpoints from EHRs, surpassing NLP methods.
- LLM accuracy approximated expert performance without requiring domain-specific knowledge.
- These prompt engineering strategies can be applied to other domains for non-expert automated data analysis.
Related Concept Videos
Mechanism of Cardiac Arrhythmias
Disturbances in Heart Rhythm
Arrhythmias are categorized by their speed, rhythm, and origin. A slow...
Pulse rhythm
Conversely, an irregular pulse pattern is termed dysrhythmia, stemming from disruptions in cardiac...

