Related Experiment Video
Updated: May 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Artificial Intelligence for Biomedical Diagnostics: Diagnostic Accuracy and Reliability of Multimodal Large Language
Henrik Stelling1,2, Armin Kraus3, Gerrit Grieb4,5
1Practices for Nuclear Medicine, Rubensstraße 125, 12157 Berlin, Germany.
Life (Basel, Switzerland)
|May 4, 2026
Summary
Multimodal large language models (MLLMs) show limited accuracy in interpreting electrocardiograms (ECGs), failing to meet clinical reliability standards for cardiovascular diagnostics. Further research is needed to improve AI performance and ensure trustworthy results.
Area of Science:
- Cardiology
- Artificial Intelligence
- Medical Diagnostics
Background:
- Electrocardiogram (ECG) interpretation is crucial for cardiovascular diagnostics but requires expertise and is prone to variability.
- Multimodal large language models (MLLMs) demonstrate potential in medical image analysis, yet their efficacy in ECG interpretation is not well-established.
Purpose of the Study:
- To evaluate the diagnostic accuracy and inter-run reliability of five leading MLLMs for standard 12-lead ECG interpretation tasks.
- To compare MLLM performance across different ECG interpretation categories and assess their clinical applicability.
Main Methods:
- Five MLLMs (ChatGPT-5.3, Gemini 3.1 Pro, Claude Opus 4.6, Grok 4.1, ERNIE 5.0) analyzed 13 standard 12-lead ECGs over five independent runs each.
- Six categorical ECG interpretation tasks and heart rate estimation were assessed against expert-consensus ground truth, with accuracy and mean absolute error (MAE) calculated.
Main Results:
- Overall categorical accuracy for MLLMs ranged from 52.3% to 64.9%.
- QRS duration classification showed the highest accuracy (66.2-90.8%), while ST/T-wave morphology assessment had the lowest performance (20.0-41.5%).
- Heart rate MAE varied between 14.8 and 46.7 bpm, with observed dissociation between accuracy and reliability.
Conclusions:
- Current MLLMs do not achieve clinically reliable performance for ECG interpretation.
- The study underscores the necessity of evaluating both diagnostic accuracy and inter-run reliability for AI systems in biomedical diagnostics.