Related Experiment Video
Updated: Jan 9, 2026

Home-Based Prescribed Pulmonary Exercise in Patients with Stable Chronic Obstructive Pulmonary Disease
Published on: August 24, 2019
ChatGPT-3.5 and the Polish thoracic surgery specialty examination: a performance evaluation
Adam Mitręga1, Dominika Kaczyńska1, Mikołaj Magiera1
1Students' Scientific Association of Computer Analysis and Artificial Intelligence at the Department of Radiology and Nuclear Medicine, Medical University of Silesia, Katowice, Poland.
ChatGPT-3.5 scored 42.2% on the thoracic surgery specialist exam, below the 60% passing threshold. This artificial intelligence tool shows limitations in clinical knowledge and consistency, requiring caution for specialized assessments.
Area of Science:
- Medical Education
- Artificial Intelligence in Medicine
- Thoracic Surgery
Background:
- Rapid advancements in artificial intelligence (AI) present new opportunities for medical applications.
- The reliability and limitations of AI in healthcare require thorough evaluation.
- Assessing AI's performance in specialized medical examinations is crucial.
Purpose of the Study:
- To evaluate the effectiveness of the ChatGPT-3.5 language model.
- To assess ChatGPT-3.5's performance on the National Specialist Examination (PES) for thoracic surgery.
Main Methods:
- Analysis of 120 test questions from the 2015 PES examination.
- Questions categorized by subject matter, clinical character, and cognitive requirements.
- ChatGPT-3.5 answered each question five times; statistical tests (χ², Kruskal-Wallis, Mann-Whitney, Spearman) and Fleiss' k coefficient were used.
Main Results:
- ChatGPT-3.5 achieved 42.2% correct answers, falling short of the 60% passing score.
- A statistically significant difference was noted between clinical and non-clinical questions (p=0.041).
- Correct answers had higher confidence (p<0.001), but confidence did not correlate with psychometric indicators; response consistency was moderate (k=0.341).
Conclusions:
- ChatGPT-3.5's performance is equivalent to a failing score on the specialist examination.
- While response confidence correlates with correctness, AI limitations in clinical knowledge and consistency necessitate caution.
- Further research is needed to understand AI's role in specialized medical knowledge assessment.
Related Concept Videos
Pulmonary Function Tests
Pulmonary Function Tests are crucial diagnostic tools for assessing respiratory function, particularly in patients with chronic respiratory disorders. They comprehensively evaluate lung volumes, ventilatory function, breathing mechanics, diffusion, and gas exchange. These tests help diagnose pulmonary diseases and play a significant role in monitoring disease progression, evaluating disability, and assessing response to therapy.
PFTs involve using a spirometer, a...
Flail Chest-II
Assessment:
1. Clinical Evaluation:
History:
Endoscopic Studies II: Thoracocentesis
Description
Excess pleural fluid or air may accumulate in some respiratory disorders in the thoracic cavity. To treat pleural effusion, a physician conducts thoracentesis by carefully piercing the chest wall and entering...
Physical Assessment of the Respiratory Tract II: Palpation
Thoracic Palpation
Thoracic palpation detects tenderness, masses, lesions, respiratory excursions, and vocal fremitus. The nurse assesses...
Physical Assessment of the Respiratory Tract III: Percussion
Percussion in Respiratory Assessment
Percussion evaluates underlying tissue composition with audible and tactile vibrations,...
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...

