Related Experiment Video
Updated: Sep 10, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Five advanced chatbots solving European Diploma in Radiology (EDiR) text-based questions: differences in performance
Jakub Pristoupil1, Laura Oleaga2, Vanesa Junquero3
1Department of Imaging Methods, Motol University Hospital and Second Faculty of Medicine, Charles University, Prague, Czech Republic.
Claude 3.5 Sonnet led in performance, confidence, and consistency for European Diploma in Radiology (EDiR) questions. All tested AI chatbots surpassed human performance on these EDiR exam questions.
Area of Science:
- Artificial Intelligence in Medical Education
- Large Language Models in Radiology
- Chatbot Performance Evaluation
Background:
- Assessing the capabilities of AI-powered chatbots in medical examinations.
- Evaluating chatbot performance, confidence, and response consistency.
Purpose of the Study:
- To compare five large language model chatbots on their performance in solving European Diploma in Radiology (EDiR) multiple-response questions.
- To determine which chatbot exhibits superior accuracy, confidence, and response consistency.
Main Methods:
- Five chatbots (ChatGPT-4o, ChatGPT-4o-mini, Copilot, Gemini, Claude 3.5 Sonnet) were tested on 52 EDiR text-based multiple-response questions.
- Chatbots provided answers, confidence scores (0-10), and were evaluated over two iterations.
- A weighted formula calculated scores per question (0.0-1.0).
Main Results:
- Claude 3.5 Sonnet achieved the highest score (0.84 ± 0.26), followed by ChatGPT-4o (0.76 ± 0.31).
- Claude 3.5 Sonnet reported the highest confidence (9.0 ± 0.9) and demonstrated superior consistency (5.4% response variation).
- All chatbots outperformed previous human EDiR candidates on these questions.
Conclusions:
- Claude 3.5 Sonnet demonstrated superior accuracy, confidence, and consistency in solving EDiR questions.
- ChatGPT-4o showed strong performance, nearly matching Claude 3.5 Sonnet.
- Significant performance variations among chatbots necessitate cautious deployment in high-stakes settings.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
06:37Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Non-equilibrium in the Cell
Machines: Problem Solving II
Machines: Problem Solving I
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
Distribution Reliability and Automation
Problem-Solving