Related Experiment Video
Updated: May 15, 2026

Retrospective Cardiac Gating with A Prototype Small-Animal X-ray Computed Tomograph
Published on: February 21, 2025
ChatGPT-4.0 and Medical Students: A Recognition-Gated Comparative Evaluation on Image-Based Medical Examinations
Andreas Sarantopoulos1, Zoi Dorothea Pana2, Andreas Larentzakis1
1School of Medicine, European University Cyprus, Nicosia, Cyprus.
Background:
Artificial intelligence (AI), particularly large language models like vision-capable ChatGPT 4.0, is increasingly shaping medical education. While these systems show promise for automated feedback and adaptive assessments, their performance in visually intensive, image-based disciplines remains insufficiently studied.
Objective:
This cohort study aims to compare the performance of ChatGPT 4.0 and undergraduate medical students on standardized, image-based multiple-choice questions in Anatomy, Pathology, and Pediatrics. Standardized exams were administered to second-, third-, and fifth-year students, and the same questions were submitted to ChatGPT 4.0 using a two-step deterministic and stochastic protocol. Items with images that ChatGPT 4.0 failed to recognize were excluded. The statistical unit of analysis was the question, and all questions were analyzed as independent question-level observations within each domain. Paired t-tests or Wilcoxon signed-rank tests were used as appropriate, and subgroup analyses were restricted to questions with a discrimination index ≥ 0.1.
Results:
Of 90 questions, 52 were eligible for the primary comparative analysis after exclusion of items in which ChatGPT 4.0 failed image recognition. ChatGPT 4.0 significantly underperformed students in Anatomy (mean difference = -0.387, p < 0.00001, Cohen's d = 2.10) but outperformed students in Pediatrics (mean difference = +0.174, p = 0.00013, Cohen's d = 0.81); these findings were similar or stronger in discrimination-based subgroup analyses. No Pathology items were eligible for comparative analysis because ChatGPT 4.0 failed image recognition for all Pathology images; therefore, no inference about comparative downstream reasoning can be made for Pathology under this protocol. In the global end-to-end analysis, which scored image-recognition failures as incorrect, ChatGPT 4.0 accuracy was 17.3% in Anatomy, 84.9% in Pediatrics, and 0% in Pathology.
Conclusion:
These findings demonstrate marked variability in ChatGPT's visual reasoning across medical domains, underscoring the need for multimodal integration and critical evaluation of AI applications before adoption in image-dependent medical education settings.
Related Concept Videos
Imaging Studies III: Computed Tomography
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Imaging Studies IV: Magnetic Resonance Imaging
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...

