大型语言模型性能的进步:对ABSITE问题上的ChatGPT-4和ChatGPT-5进行比较研究
Ahmet Necati Sanli1, Bilal Turan2, Deniz Esin Tekcan Sanli3
1Department of General Surgery, Abdulkadir Yuksel State Hospital, Gaziantep, Turkey.
The American surgeon
|October 18, 2025
概括
聊天GPT-5在美国外科医学会培训考试 (ABSITE) 测试中显著超过了聊天GPT-4o,在案例场景中显示出更好的准确性. 这表明新的AI模型可以更好地支持外科教育和考试准备.
科学领域:
- 医疗教育中的人工智能
- 手术培训技术手术培训技术
- 大型语言模型 (LLM)
背景情况:
- 美国外科医学会实习生考试 (ABSITE) 是对外科实习生进行的关键性评估.
- 评估像ChatGPT这样的先进AI模型在医学知识回忆和应用中的性能是必不可少的.
- 后续的AI版本,如ChatGPT-4o和ChatGPT-5,可能在复杂的问题回答方面提供不同的功能.
研究的目的:
- 直接比较ChatGPT-4o和ChatGPT-5在ABSITE问卷内容上的问答性能.
- 在ABSITE测试中确定AI模型之间的性能差异最明显的特定领域.
- 评估高级LLM在帮助外科教育和考试准备方面的潜力.
主要方法:
- 使用了170个多选项ABSITE测试题,从2017年到2022年.
- 问题分为定义,生物化学/制药,病例情景和治疗和手术程序.
- 聊天GPT-4o和聊天GPT-5都回答了相同的问题,使用麦克纳马的测试记录和分析准确率.
主要成果:
- 聊天GPT-5的整体准确率为87.1%,而聊天GPT-4o的整体准确率为79.4% (P < 0.001).
- 在案例场景问题 (76.3%至86.8%,P=0.008) 中观察到显著改善,这表明临床推理更好.
- 由于上限效应,在定义或生物化学/制药类别中没有发现显著差异;治疗和外科手术程序显示不显著的改善.
结论:
- 在ABSITE测试问题上,ChatGPT-5表现优于ChatGPT-4o,特别是在复杂的基于案例的场景中.
- 这些发现表明,较新的LLM版本为准备考试的外科实习生提供了更可靠的帮助.
- 在现实世界考试环境和多式联运应用中进行进一步的验证是有必要的,以确认这些好处.
相关概念视频
Comparing Experimental Results: Student's t-Test
4.9K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.9K
Multiple Comparison Tests
4.4K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.4K


