Related Experiment Videos
Comparison of emergency physicians and artificial intelligence models in pneumothorax detection: A multi-reader
Omer Faruk Cakiroglu1, Bilal Arac2, Seyma Nur Calisir3
1Department of Emergency Medicine, Kartal Doktor Lutfi Kirdar City Hospital, Istanbul, Turkiye.
Medicine
|June 23, 2026
Summary
General-purpose AI models like ChatGPT and Gemini show varied performance in detecting pneumothorax on chest X-rays. While physicians remain superior, these AI tools may offer adjunctive support after further optimization.
Area of Science:
- Radiology
- Artificial Intelligence
- Medical Diagnostics
Background:
- Pneumothorax (PTX) detection on chest radiographs (CXRs) is crucial for timely patient management.
- General-purpose large language models (LLMs) are increasingly explored for medical image analysis.
- Evaluating LLM performance against human experts is essential for clinical integration.
Purpose of the Study:
- To compare the diagnostic performance of ChatGPT and Gemini for pneumothorax detection on CXRs.
- To assess the performance of these LLMs against emergency physicians.
- To analyze the impact of case difficulty on LLM diagnostic accuracy.
Main Methods:
- Retrospective analysis of 265 PTX and 267 non-PTX adult CXR cases.
- Independent, blinded review of CXRs by 13 emergency physicians.
- Evaluation of the same CXRs by ChatGPT and Gemini using standardized prompts.
- Comparison of diagnostic metrics including sensitivity, specificity, accuracy, and kappa agreement.
Main Results:
- ChatGPT: 44.5% sensitivity, 95.5% specificity, 70.1% accuracy (kappa=0.401).
- Gemini: 52.5% sensitivity, 79.0% specificity, 65.8% accuracy (kappa=0.315).
- Emergency physicians achieved 64.5% sensitivity, 99.6% specificity, and 82.1% accuracy.
- Both LLMs showed decreased accuracy with increasing case difficulty (ChatGPT r=-0.438, Gemini r=-0.274).
Conclusions:
- General-purpose LLMs exhibit distinct performance profiles for PTX detection, with model-specific strengths and weaknesses.
- Physician performance in PTX detection on CXRs remains superior to current general-purpose LLMs.
- LLMs may serve as adjunctive tools in radiology but require task-specific optimization and clinical validation before widespread adoption.