Related Experiment Video
Updated: Sep 18, 2026

Porcine Cardiac Arrest Model Using an Implantable Defibrillator
Published on: January 9, 2026
Large language models fail to reliably predict emergent catheterization laboratory activation from prehospital
Emile Legendre1, Urska Cvek2, Stewart Greathouse3
1Louisiana State University Health Sciences Center-Shreveport, 1501 Kings Highway, Shreveport, LA, 71103, USA.
Background:
Rapid and accurate electrocardiogram (ECG) interpretation is essential for timely identification of ST-elevation myocardial infarction (STEMI) and activation of reperfusion pathways in emergency care.
Objectives:
To evaluate the diagnostic performance of multimodal LLMs in identifying prehospital ECGs warranting emergent catheterization laboratory activation.
Methods:
We performed a retrospective analysis of 615 ECGs from 270 emergency medical service patient encounters (EMS) with concern for acute myocardial infarction. The reference standard was cardiology activation of the STEMI pathway for emergent angiography. LLM-based image interpretation (three models) and ECG machine algorithm interpretations were compared. Sensitivity, specificity, positive predictive value, negative predictive value, and overall accuracy were calculated.
Results:
Gemini demonstrated the highest sensitivity (95.3%; 95% CI 91.7-97.3) but extremely poor specificity (9.4%), indicating a high false-positive rate. ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%). The ECG machine algorithm demonstrated more balanced performance, with sensitivity of 67.7% (95% CI 61.4-73.4) and higher specificity (64.2%) than all LLMs.
Conclusions:
Multimodal LLM interpretation of prehospital ECGs demonstrated clinically unreliable performance for identifying ECGs warranting emergent cardiac catheterization laboratory activation when benchmarked against real-world cardiology activation decisions. Although some models achieved high sensitivity, poor specificity resulted in excessive false-positive activation recommendations. These findings suggest that general-purpose LLMs are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive cardiopulmonary care workflows.