在兽医招生中,机器等级 (ChatGPT) 和人类等级作文分数的比较
Raphael Vanderstichel1, Henrik Stryhn2
1Department of Veterinary Clinical Sciences, College of Veterinary Medicine, Long Island University, Brookville, NY, 11548, USA.
Journal of veterinary medical education
|November 6, 2024
概括
人工智能作文评分显示了比人类评分器更高的精度,与人类得分相对应适度,与认知措施稍好一些. 谨慎的快速设计对于公平的基于机器的入学审查至关重要.
科学领域:
- 医学教育 医学教育
- 教育中的人工智能
- 招生程序 招生程序 招生程序
背景情况:
- 传统的入院依赖于认知测量.
- 整体审查现在通过散文,采访和情境判断测试强调非认知技能.
- 整体审查需要大量的教师资源.
研究的目的:
- 为了评估人与机器之间的协议 (OpenAI的ChatGPT) 文章评分.
- 计算论文分数与其他招生标准之间的相关性.
- 为了评估机器与人类作文评分的精度.
主要方法:
- 使用自然语言处理 (NLP) 来进行机器分级.
- 人类和机器作文得分使用相关性分析进行了比较.
- 对于这两种评分方法来说,计算了inter-rater和inter-replicate可靠性.
主要成果:
- 机器分级显示了更高的相互复制可靠性 (0.410.61) 比人类相互评分可靠性 (0.070.41).
- 在人类和机器作文得分之间发现了中等相关性 (0.41).
- 机器分数与人类分数相比,与认知指标的相关性略有强烈,精度更高 (2-3倍).
结论:
- 人工智能展示了准确可靠的入学作文评分的潜力.
- 精心及时的工程和标签开发对于机器分级至关重要.
- 统计分析和复制对于使用人工智能进行公平的申请人评估至关重要.
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Ethics in Research
22.9K
Today, scientists agree that good research is ethical in nature and is guided by a basic respect for human dignity and safety. However, this has not always been the case. Modern researchers must demonstrate that the research they perform is ethically sound.
22.9K


