Related Experiment Video
Updated: Jan 21, 2026

A Porcine Model of Acute Autologous Pulmonary Embolism
Published on: September 6, 2024
AI in Patient Care: Evaluating Large Language Model Performance Against Evidence-Based Guidelines for Pulmonary
Ömer F Karakoyun1, Halil E Koyuncuoğlu1, Ömer H Sağnıç2
1Clinic of Emergency Medicine, Muğla Training and Research Hospital, Muğla, Türkiye.
Four artificial intelligence (AI) large language models (LLMs) were evaluated for their ability to apply pulmonary embolism (PE) clinical guidelines. ChatGPT-4o performed best, but further development is needed for consistent guideline adherence in physician workflows.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Decision Support
Background:
- Artificial intelligence (AI)-driven large language models (LLMs) are increasingly integrated into healthcare for patient education.
- The efficacy of LLMs in interpreting and applying complex clinical guidelines within real-world physician workflows remains largely unassessed.
- Pulmonary embolism (PE), with its defined management protocols, serves as an ideal model for evaluating LLM performance in clinical guideline application.
Purpose of the Study:
- To assess the performance of four leading AI-driven LLMs (ChatGPT-4o, DeepSeek-V2, Gemini, Grok) in applying the 2019 European Society of Cardiology guidelines for PE.
- To evaluate the clinical accuracy, guideline adherence, and response consistency of these LLMs when presented with a simulated PE case.
Main Methods:
- Ten open-ended questions based on a simulated PE case were developed, covering diagnosis, risk stratification, treatment, and follow-up.
- LLM responses were scored by two emergency physicians using a 10-point scale against guideline-based reference answers.
- Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC), and group comparisons were made using Kruskal-Wallis tests.
Main Results:
- ChatGPT-4o achieved the highest overall score (76), followed by Gemini (73.75), Grok (71.25), and DeepSeek-V2 (65), with no statistically significant difference in total scores (P = 0.390).
- Performance varied across categories, with ChatGPT-4o excelling in follow-up and DeepSeek-V2 in diagnostics.
- Expert reviewers noted strengths like ChatGPT-4o's structured responses and Grok's practicality, while identifying limitations such as insufficient personalization and guideline gaps. Inter-rater agreement was excellent (ICC: 0.986).
Conclusions:
- AI-driven LLMs demonstrate potential utility in supporting the management of pulmonary embolism.
- No single LLM consistently outperformed others across all assessed domains of PE management.
- Further advancements are necessary to improve the clinical integration and consistent guideline compliance of LLMs in healthcare settings.
More Related Videos
08:02Establishment of a Minimally Invasive Rat Model of Pulmonary Embolism Using Autologous Blood Clots
Published on: October 25, 2024
06:15Protocol and Guidelines for Point-of-Care Lung Ultrasound in Diagnosing Neonatal Pulmonary Diseases Based on International Expert Consensus
Published on: March 6, 2019
Related Concept Videos
Pulmonary Embolism II: Diagnostic Studies and Interprofessional Care
Pulmonary Embolism I: Introduction
Pulmonary Embolism III: Nursing Management
Patient-centered Care
The Evidence for Evolution
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...