Related Experiment Video
Updated: Jul 15, 2025

05:04
Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training
Published on: August 9, 2024
978
Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments
Dana Brin1,2, Vera Sorin3,4, Akhil Vaid5
1Department of Diagnostic Imaging, Chaim Sheba Medical Center, Ramat Gan, Israel. dannabrin@gmail.com.
Scientific Reports
|October 1, 2023
Summary
Artificial intelligence (AI) models like GPT-4 show promise in answering United States Medical Licensing Examination (USMLE) questions on medical soft skills, outperforming previous AI and human users in communication, ethics, and empathy.
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
- Medical Ethics
Background:
- The United States Medical Licensing Examination (USMLE) assesses physician competency.
- AI models have been studied for USMLE performance, but soft skills assessment is lacking.
- Evaluating AI in communication, ethics, empathy, and professionalism is crucial for medical practice.
Purpose of the Study:
- To assess the performance of ChatGPT and GPT-4 on USMLE-style soft skills questions.
- To compare AI models' capabilities against past human performance on the AMBOSS platform.
- To evaluate AI consistency and confidence in responding to complex medical scenarios.
Main Methods:
- Utilized 80 USMLE-style soft skills questions from official sources and AMBOSS.
- Administered questions to ChatGPT and GPT-4, employing follow-up queries for consistency checks.
- Compared AI model performance against historical data from AMBOSS users.
Main Results:
- GPT-4 achieved 90% accuracy, significantly outperforming ChatGPT's 62.5%.
- GPT-4 demonstrated higher confidence, with no response revisions, unlike ChatGPT (82.5% revisions).
- GPT-4's performance surpassed that of previous AMBOSS users, showing empathy and professionalism.
Conclusions:
- GPT-4 exhibits strong potential for handling USMLE soft skills questions, exceeding both earlier AI and human benchmarks.
- AI models, particularly GPT-4, demonstrate capacity for empathy and ethical reasoning in medical contexts.
- AI may offer valuable support in addressing the interpersonal and ethical demands of medical practice.
More Related Videos
Related Concept Videos
Comparing Experimental Results: Student's t-Test
1.6K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
1.6K
Comparing the Survival Analysis of Two or More Groups
218
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
218

