在USMLE软技能评估中比较ChatGPT和GPT-4的表现
Dana Brin1,2, Vera Sorin3,4, Akhil Vaid5
1Department of Diagnostic Imaging, Chaim Sheba Medical Center, Ramat Gan, Israel. dannabrin@gmail.com.
Scientific reports
|October 1, 2023
概括
像GPT-4这样的人工智能 (AI) 模型在回答美国医疗执照考试 (USMLE) 关于医疗软技能的问题方面表现有希望,在沟通,道德和同理心方面表现优于以前的AI和人类用户.
科学领域:
- 医学教育 医学教育
- 医疗保健中的人工智能
- 医学伦理 医学伦理
背景情况:
- 美国医学执照考试 (USMLE) 评估医生的能力.
- 人工智能模型已经研究了USMLE绩效,但软技能评估缺乏.
- 在医疗实践中,评估人工智能在沟通,道德,同情和专业精神方面至关重要.
研究的目的:
- 评估ChatGPT和GPT-4在USMLE风格软技能问题上的表现.
- 将AI模型的能力与AMBOSS平台上人类过去的表现进行比较.
- 评估AI的一致性和对复杂医疗场景的反应信心.
主要方法:
- 使用了80个USMLE风格的软技能问题,来自官方来源和AMBOSS.
- 向ChatGPT和GPT-4管理问题,使用后续查询进行一致性检查.
- 将AI模型性能与AMBOSS用户的历史数据进行比较.
主要成果:
- GPT-4实现了90%的准确性,明显超过了ChatGPT的62.5%的性能.
- 与ChatGPT (82.5%的修订) 不同的是,GPT-4表现出更高的信心,没有响应修订.
- GPT-4的表现超过了之前的AMBOSS用户的表现,表现出同情心和专业精神.
结论:
- 在处理USMLE软技能问题方面,GPT-4具有很强的潜力,超过了先前的人工智能和人类基准.
- 人工智能模型,特别是GPT-4,在医疗环境中展示了同情心和道德推理的能力.
- 人工智能可以在解决医疗实践的人际关系和道德要求方面提供有价值的支持.
相关概念视频
Comparing Experimental Results: Student's t-Test
1.6K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
1.6K
Comparing the Survival Analysis of Two or More Groups
218
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
218


