Related Experiment Videos
Evaluating large language models for myocardial infarction public health education: a comparative study on
Tailong Lv1, Wenkai Bao2, Shudi Li1
1Shandong University of Traditional Chinese Medicine, Jinan, China.
Objective:
Myocardial infarction (MI) is an acute, life-threatening cardiovascular disease, and high-quality, accessible public health education is vital for emergency management. This study systematically evaluates the quality, transparency, clinical accuracy, patient safety, and readability of information generated by different large language models (LLMs) in responding to MI-related public inquiries.
Methods:
Twenty-five representative MI patient education questions were submitted to Gemini 3.5 Flash, Claude Opus 4.8, and ChatGPT 5.5. The generated information was independently evaluated by two cardiologists using four validated tools (DISCERN, EQIP, GQS, and JAMA) alongside a strict clinical safety assessment. Text readability was concurrently assessed utilizing six established metrics (FRES, ARI, GFI, CLI, FKGL, and SMOG).
Results:
Significant variations were observed in the quality, transparency and readability of information generated by the evaluated LLMs. Regarding quality and transparency, significant overall differences were noted among models in DISCERN (p < 0.001) and EQIP (p < 0.001) scores, whereas no significant differences were found in GQS and JAMA benchmarks. For DISCERN, Claude achieved significantly higher scores than both Gemini and ChatGPT. In the EQIP assessment of completeness and clarity, Claude and Gemini scored significantly higher than ChatGPT, though all models attained a "good" rating. Additionally, all models exhibited poor performance on the JAMA benchmark, indicating critical deficits in information transparency. Notably, LLMs sometimes generated incomplete, incorrect, or even potentially harmful information. Regarding readability, although Claude generated relatively more comprehensible text, all models failed to meet the recommended sixth-grade reading benchmark, indicating high reading difficulty.
Conclusion:
While LLMs can generate structurally clear and logically coherent foundational content for MI-related queries, they occasionally produce clinically inappropriate directives. Furthermore, the texts generated by these models are overly complex, creating substantial reading barriers for the general public. Consequently, under zero-shot and English-language testing conditions, the current LLMs are not yet capable as standalone health education tools for MI.