Related Experiment Video
Updated: Apr 25, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Cardiac MRI Diagnosis Based on Standardized Text Descriptions
Hongbo Zhang1, Junjie Zhou2, Chen Zhang1
1Department of Interventional Diagnosis and Treatment, Beijing Anzhen Hospital Affiliated to Capital Medical University, Beijing, China.
Background:
MRI is important for cardiac disease evaluation, but accurate diagnosis remains challenging in less experienced centers. Although large language models (LLMs) have shown promise in medical imaging diagnosis, their application in cardiac MRI is limited.
Hypothesis:
LLMs may be effective in achieving cardiac MRI diagnosis based on standardized descriptions.
Study Type:
Retrospective.
Population:
A total of 203 hypertrophic cardiomyopathy, 186 dilated cardiomyopathy, 46 hypertensive heart disease, 198 ischemic cardiomyopathy, 38 constrictive pericarditis, 45 cardiac amyloidosis, 91 myocarditis, and 144 normal controls.
Field Strength/Sequences:
Balanced steady-state free-precession, short tau inversion recovery, and breath-hold inversion-recovery segmented gradient-echo sequences at 3.0 T.
Assessment:
Clinical and cardiac MRI information from each subject was converted into standardized descriptions and input into Generative Pre-trained Transformer-4.5 (GPT-4.5), GPT-4 Omni (GPT-4o), Deepseek-V3, and Deepseek-R1 LLMs. Cardiac MRI information included LV function, wall thickness and motion, and abnormalities on T2WI, perfusion and late gadolinium enhancement sequences. Each model was asked to generate an imaging diagnosis. In addition, a medical student (8 months experience) and three radiologists (junior, mid-level and senior: with 3, 6, and 10 years' experience, respectively) provided diagnoses based on cardiac MRI images and clinical information.
Statistic Tests:
Frequency-weighted sensitivity and specificity were calculated. The diagnostic performances of the LLMs and human readers were compared using the McNemar test with Bonferroni correction. A p value < 0.05 was considered significant.
Results:
All LLMs showed excellent frequency-weighted specificity (0.973-0.983). The frequency-weighted sensitivities of all LLMs were not significantly different from that of the junior radiologist, were significantly higher than that of the medical student, and significantly inferior to those of the senior radiologist (GPT-4.5: 0.863, GPT-4o: 0.821, Deepseek-V3: 0.843, and Deepseek-R1: 0.851 vs. junior radiologist: 0.850, all adjusted p = 1.000; vs. medical student: 0.731, all adjusted p < 0.001; vs. senior radiologist: 0.942, all adjusted p < 0.001). Additionally, the mid-level radiologist achieved a frequency-weighted sensitivity of 0.895, outperforming all LLMs except GPT-4.5.
Data Conclusion:
LLMs may generate accurate diagnoses from standardized cardiac MRI descriptions, potentially benefiting less experienced physicians.
Technical Efficacy:
Stage 5.
More Related Videos
12:15Tissue Preparation Techniques for Contrast-Enhanced Micro Computed Tomography Imaging of Large Mammalian Cardiac Models with Chronic Disease
Published on: February 8, 2022
09:57Development and Evaluation of 3D-Printed Cardiovascular Phantoms for Interventional Planning and Training
Published on: January 18, 2021