当人工智能模型参加考试时:大型语言模型与医学学生在多项选择课程考试中
Pablo Ros-Arlanzón1,2, Renato Gutarra-Ávila3, Vicente Arrarte-Esteban3,4
1Neurology Department, Dr. Balmis General University Hospital, Alicante, Spain.
Medical education online
|November 29, 2025
概括
大型语言模型 (LLM) 在西班牙语多选择题考试中明显优于医学学生,即使是负分. LLM 显示出高精度和可重复性,这表明在医学教育评估中有监督使用的潜力.
科学领域:
- 医学教育 医学教育
- 医疗保健中的人工智能
- 评估方法 评估方法
背景情况:
- 大型语言模型 (LLM) 越来越多地融入医疗保健和医学教育.
- 在机构授权的多项选择题 (MCQ) 上,LLM的表现,特别是负分的表现,并未得到充分理解.
- 在标准化医疗评估中评估LLM能力对于其负责任的实施至关重要.
研究的目的:
- 为了比较五个当代LLM的考试成绩,与在最终课程MCQ考试中注册的医学学生进行比较.
- 评估LLM准确性和可重复性在西班牙语临床课程考试负分标记.
- 探索LLM在医学MCQ评估中的潜在作用.
主要方法:
- 在米格尔·赫尔南德斯大学 (Miguel Hernández University) 进行了一项比较的横截面研究,使用四个临床课程 (传染病,神经病学,呼吸系统医学,心血管医学) 的结论考试.
- 五个LLM (OpenAI o1,GPT-4o,DeepSeek R1,Microsoft Copilot,Google Gemini 1.5 Flash) 在两个独立的运行中完成了所有MCQ.
- 分析了学生成绩 (n=442) 和法学士成绩,并使用Gwet的AC1.1估计了测试重复测试的可靠性.
主要成果:
- 在所有课程中,LLM的得分始终高于学生的平均成绩,平均得分从7.46-9.88不等,而学生平均得分为4.28-7.32.
- 在三个课程中,OpenAI o1获得了最高的平均分数,而Copilot在心血管医学 (仅文本子集) 中领先.
- 所有LLM都回答了每一个MCQ,并且观察到高短期测试重试协议 (AC1 0.79-1.00).
结论:
- 在西班牙的MCQ考试中,LLM与医学学生相比表现优越,并获得负分.
- 法律法规表现出高准确性和短期可重复性,支持其作为监督评估辅助的潜在用途.
- 需要进一步的研究来证实这些发现在不同的机构,语言和问题格式,并评估他们的教育影响.
关键词:
人工智能的人工智能聊天GPT 聊天 在GPT 聊天大型语言模型.副飞行员是第二名的.在深度搜索中,深度搜索.双子座 (Gemini) 是一个双子座.医学教育 医学教育医学院的学生 医学院的学生多选题的问题是多选题.更多相关视频
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
8.0K
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
相关概念视频
Multiple Comparison Tests
4.4K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.4K
Comparing Experimental Results: Student's t-Test
4.7K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.7K
