Related Experiment Video
Updated: Jan 11, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
AI performance in emergency medicine fellowship examination: comparative analysis of ChatGPT-4o, Gemini 2.0, Claude
İshak Şan1, Medine Akkan Öz2, Mehmet Yortanli3
1Department of Emergency Medicine, Health Science University, Bilkent City Hospital, Ankara, Turkiye.
Background/Aim:
This study evaluated the accuracy rates and response consistency of four different large language models (ChatGPT-4o, Gemini 2.0, Claude 3.5, and DeepSeek R1) in answering questions from the Emergency Medicine Fellowship Examination (YDUS), which was administered for the first time in Türkiye.
Materials And Methods:
In this observational study, 60 multiple-choice questions from the Emergency Medicine YDUS administered on 15 December 2024, were classified as knowledge-based (n = 26), visual content (n = 2), and case-based (n = 32). Each question was presented three times to the four large language models. The models' accuracy rates were evaluated according to overall accuracy, strict accuracy, and ideal accuracy criteria. Response consistency was measured using Fleiss' Kappa test.
Results:
The ChatGPT-4o model was the most successful in terms of overall accuracy (90.0%), while DeepSeek R1 showed the lowest performance (76.7%). Claude 3.5 (83.3%) and Gemini 2.0 (80.0%) demonstrated moderate success. When analyzed by category, ChatGPT-4o achieved the highest success with 92.3% accuracy in knowledge-based questions and 90.6% in case-based questions. In terms of response consistency, the Claude 3.5 model (Fleiss' Kappa = 0.68) showed the highest consistency, while Gemini 2.0 (Fleiss' Kappa = 0.49) showed the lowest. Inconsistent hallucinations were more frequent in the Gemini 2.0 and DeepSeek R1 models, whereas persistent hallucinations were less common in the ChatGPT-4o and Claude 3.5 models.
Conclusion:
Large language models can achieve high accuracy rates for knowledge and clinical reasoning questions in emergency medicine but show differences in terms of response consistency and hallucination tendency. While these models have significant potential for use in medical education and as clinical decision support systems (CDSS), they need further development to provide reliable, up-to-date, and accurate information.
More Related Videos
06:19Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024