使用RaschOnline评估ChatGPT对多选题的能力:观察性研究
Julie Chi Chow1,2, Teng Yun Cheng3, Tsair-Wei Chien4
1Department of Pediatrics, Chi Mei Medical Center, Tainan, Taiwan.
JMIR formative research
|August 8, 2024
概括
在2023年台湾大学入学考试中,ChatGPT在回答多项选择题 (MCQ) 方面表现出"A"级的熟练程度. 这项研究使用了Rasch分析 (RaschOnline) 来评估AI.
科学领域:
- 人工智能的人工智能
- 教育测量教育的测量
- 心理测量 心理测量 心理测量
背景情况:
- 一个领先的大型语言模型ChatGPT在专业应用中表现有前途.
- 有限的研究存在于AI在使用Rasch分析的多选择题 (MCQ) 的表现.
- 拉什分析中的KIDMAP是评估AI的MCQ回答能力的工具.
研究的目的:
- 为了证明RaschOnline对评估AI性能的实用性.
- 为了评估ChatGPT在MCQ上的表现,与正常样本进行对比.
- 为了确定通过ChatGPT获得的学术成绩.
主要方法:
- 分析了ChatGPT对2023年台湾大学入学考试中的10个MCQ的答案.
- 300名模拟学生使用拉什模型生成与ChatGPT进行比较.
- RaschOnline被用来生成视觉展示,包括项目难度,DIF,ICC,赖特地图和KIDMAP.
主要成果:
- 项目难度显示单调增加,逻辑从 -2.43到2.47不等.
- 在第5项中,性别组之间的差异性项目功能 (DIF) 已被注意到 (P=.04).
- 聊天GPT获得了"A"级,在B级到E级的模拟学生中表现优于模拟学生.
结论:
- RaschOnline有效地评估了在MCQ回答中的AI性能.
- 聊天GPT在回答标准化测试中的英语MCQ方面表现出卓越的熟练程度.
- 这项研究证实了ChatGPT在与人类绩效相比时能够达到高学业成绩的能力.
关键词:
聊天GPT 聊天GPT 聊天这就是KIDMAP.在RaschOnline上可以找到.赖特地图 赖特地图 赖特地图申请申请表 申请表 申请表人工智能的人工智能是人工智能.学院学院学院学院学院学院学院学院学院学院学院差异性项目的功能.评估工具是一个评估工具.多选题问题 多选题问题获得得分的得分.学生是学生的学生是学生的测试 测试 测试 测试 测试工具 工具 工具 工具网站工具 网站工具更多相关视频
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
733
10:26Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
3.9K
相关概念视频
Cochran's Q Test
276
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
276
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
Testing a Claim about Population Proportion
3.3K
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
3.3K
