Related Experiment Video
Updated: Sep 16, 2025

Home-Based Prescribed Pulmonary Exercise in Patients with Stable Chronic Obstructive Pulmonary Disease
Published on: August 24, 2019
ChatGPT Performance Deteriorated in Patients with Comorbidities When Providing Cardiological Therapeutic
Wen-Rui Hao1,2,3, Chun-Chao Chen1,2,3, Kuan Chen4
1Taipei Heart Institute, Taipei Medical University, Taipei 11002, Taiwan.
Large language models (LLMs) show promise in medication advice for complex cardiovascular disease (CVD) patients, but exhibit inconsistent performance and safety risks. Professional oversight is crucial before autonomous clinical use.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Cardiovascular Pharmacology
Background:
- Large language models (LLMs) are being explored for medical applications.
- Reliability of LLMs for medication advice in complex patient cases, especially with comorbidities, is under-investigated.
- This study addresses the need to evaluate LLM performance in complex cardiovascular disease (CVD) scenarios.
Purpose of the Study:
- To systematically evaluate the performance, consistency, and safety of ChatGPT (GPT-3.5 and GPT-4.0) in generating medication recommendations for complex CVD cases.
- To compare the accuracy and reliability of different LLM versions in simulated clinical scenarios.
- To identify potential risks and limitations of using LLMs for clinical decision support in cardiology.
Main Methods:
- A simulation-based study using 25 complex CVD scenarios, each prompted 10 times for GPT-3.5 and GPT-4.0.
- Recommendations were classified by cardiologists as "high priority" or "low priority".
- Evaluated physician approval rates, recommendation priority, response consistency (Jaccard index), and error patterns using statistical tests.
Main Results:
- GPT-4.0 showed a modestly higher overall physician approval rate (86.90%) than GPT-3.5 (85.06%), though not systematically significant for high-priority recommendations.
- Significant differences in error patterns were observed (p < 0.001), with GPT-4.0 more frequently recommending contraindicated drugs in high-risk scenarios.
- Low inter-model consistency (mean Jaccard index = 0.42) indicated frequent variations in advice between LLM responses.
Conclusions:
- Current LLMs demonstrate high overall physician approval but exhibit inconsistent performance and safety risks for complex CVD medication advice.
- LLM reliability does not yet meet standards for autonomous clinical application.
- Future research should focus on real-world data validation and domain-specific fine-tuning, emphasizing the continued need for professional oversight.
Related Concept Videos
Heart Failure VI: Adjunct Therapies
Cardiomyopathy V: Interprofessional Care
Coronary Artery Disease V: Interprofessional Care
Cardiomyopathy III: Hypertrophic Cardiomyopathy
Cardiomyopathy II: Dilated Cardiomyopathy
Angina III: Clinical Manifestations and Assessment

