聊天GPT 3.5 无法编写适当的多选项练习考试问题
Alexander Ngo1, Saumya Gupta1, Oliver Perrine1
1Department of Pathology & Laboratory Medicine, Boston University Chobanian and Avidesian School of Medicine, Boston MA, USA.
Academic pathology
|January 1, 2024
概括
人工智能 (AI) 可以影响学术教学. 虽然ChatGPT显示了创建多选择题等教育内容的潜力,但由于准确性问题,它目前需要大量的教师编辑.
科学领域:
- 教育技术的教育技术
- 教育中的人工智能
背景情况:
- 人工智能 (AI) 对传统的学术教学提出了挑战和机遇.
- 人们对像ChatGPT这样的AI工具产生原创文章存在担忧.
- 人工智能可以作为一种有价值的工具来增强现有的教学方法.
研究的目的:
- 评估ChatGPT 3.5在生成多选题 (MCQ) 及学术目的解释方面的有效性.
- 评估人工智能生成的MCQ及其解释的准确性和完整性.
主要方法:
- 利用ChatGPT 3.5根据上传的作者撰写的文本生成60个MCQ.
- 指示AI为每个MCQ提供正确和不正确答案的解释.
- 分析生成的MCQ,以确定问题的准确性,答案的正确性和解释的质量.
主要成果:
- 聊天GPT 3.5成功生成了准确的问题和答案,并解释了只有32%的MCQ (60分之19).
- 很大一部分问题 (25%) 包含错误或误导性的答案.
- 在许多情况下,ChatGPT未能为错误的答案选择提供解释.
结论:
- 像ChatGPT 3.5这样的当前人工智能模型在自主生成准确的教育评估方面显示出有限的可靠性.
- 在使用人工智能生成的内容进行实践考试或评估时,广泛的人类审查和编辑是必不可少的.
- 尽管有局限性,但人工智能工具仍然可以为教练在创建评估材料草案时提供额外的好处.
更多相关视频
08:13Development and Implementation of a Multi-Disciplinary Technology Enhanced Care Pathway for Youth and Adults with Concussion
Published on: January 20, 2019
6.6K
00:08A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
7.0K
相关概念视频
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Testing a Claim about Population Proportion
3.3K
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
3.3K
Surveys
14.8K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.8K
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Cochran's Q Test
347
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
347
