在本科医学教育中基于LLM的自动简短答案评分
1Department of Life Sciences and Medicine, University of Luxembourg, 6, avenue de la Fonte, L-4364, Esch-sur-Alzette, Luxembourg. christian.grevisse@uni.lu.
BMC medical education
|September 28, 2024
概括
大型语言模型 (LLM) 在医学教育中显示出自动简短答案分级 (ASAG) 的前景. 双子 1.0 专业版 双子 1.0 专业版
科学领域:
- 医学教育 医学教育
- 人工智能的人工智能
- 自然语言处理自然语言处理.
背景情况:
- 医学教育中传统的多选择题更倾向于承认而不是回忆.
- 对教育工作者来说,对开放式问题的评分是耗时的.
- 自动简短答案分级 (ASAG) 提供了一个解决方案,最近的大型语言模型 (LLM) 的进步推动了进展.
研究的目的:
- 评估LLM的有效性,特别是GPT-4和Gemini 1.0 Pro,用于本科医学教育中的简短答案自动评分.
- 为了比较LLM评分表现与人类评价者.
主要方法:
- 来自12个本科医学课程的2288名学生的答案分为3种语言.
- 使用GPT-4和Gemini 1.0 Pro进行分级过程.
- 研究人员将LLM成绩与人类评估员的成绩进行了比较.
主要成果:
- 双子座1.0 Pro的成绩与人类评估者非常相匹配,而GPT-4的成绩较低,但错误阳性较少.
- 两位LLM都表现出与人类等级的中度一致性,以及高精度的完全正确答案与GPT-4.
- 在高质量的答案密钥中观察到LLM评分的一致性,与答案长度或语言的相关性较弱.
结论:
- 基于LLM的ASAG在医学教育中需要人类监督,但可以节省教育工作者的时间来评分边缘病例.
- 法学士的内在知识似乎足以用于本科水平的医学教育,否定了微调的需要.
- 通过自动化评分,LLM可以帮助教育工作者,让他们更专注于复杂的学生反应.
更多相关视频
07:32Use of Galvanic Skin Responses, Salivary Biomarkers, and Self-reports to Assess Undergraduate Student Performance During a Laboratory Exam Activity
Published on: February 10, 2016
9.3K
04:12Mixed Reality for Education MRE Implementation and Results in Online Classes for Engineering
Published on: June 23, 2023
599
相关概念视频
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
