将大型语言模型的评分一致性与医学教育中形成性评估的教师进行比较
Radhika Sreedhar1, Linda Chang2, Ananya Gangopadhyaya2
1University of Illinois College of Medicine, Chicago, IL, USA. sreedhar@uic.edu.
Journal of general internal medicine
|October 14, 2024
概括
大型语言模型 (LLM) 在评估医学学生的批判性评估任务方面表现有希望,提供一致的反并减少教师的时间. 这项研究发现,LLM与教师评分相当,突出了它们在医学教育中的潜力.
科学领域:
- 医学教育 医学教育
- 医疗保健中的人工智能
- 教育技术的教育技术
背景情况:
- 医学教育需要对临床前学生的自主学习技能的个性化反.
- 使用批判性评估任务,但教师反是时间密集的.
- 大型语言模型 (LLM) 提供了自动评分和反生成的潜力.
研究的目的:
- 评估使用LLM来评估和反本科医学学生形成性评估的一致性和可行性.
- 探索基于LLM的评估的心理测量特征.
主要方法:
- 对111个临床前医学学生的临界评估作业进行横截面研究.
- 使用开发的提示符,比较ChatGPT 3.5得分与现有的教师成绩.
- 分析分数差异,评分器间可靠性 (IRR),内部一致性可靠性,精度回忆曲线下的面积 (AUCPR) 和成本效益.
主要成果:
- 单个项目的LLM评分与教师评分相当.
- 法学士和教师之间的总体一致性为67% (P < 0.001).
- 法学士使用减少了五倍的教师时间,可能节省150个教师小时.
结论:
- 在医学教育方面,LLM显示出其作为帮助教师评估培养任务和提供反的工具的潜力.
- 该研究强调了使用LLM用于教育评估的心理测量有效性和可行性.
关键词:
聊天GPT评分一致性形成性任务更多相关视频
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
7.5K
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
714
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Data Validation
5.0K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.0K
