聊天GPT的质量:概念库存项的可靠性和有效性
Stefan Küchemann1, Martina Rau2, Albrecht Schmidt3
1Chair of Physics Education Research, Faculty of Physics, Ludwig-Maximilians-Universität München (LMU Munich), Munich, Germany.
Frontiers in psychology
|October 23, 2024
概括
大型语言模型 (LLM) 可以生成物理概念项目,但质量需要仔细的快速工程和专家审查. 人类监督对于使用人工智能生成的内容进行有效的教育评估至关重要.
科学领域:
- 物理教育研究 物理学教育研究
- 教育中的人工智能
背景情况:
- 大型语言模型 (LLM) 在教育中提出了机遇和挑战.
- 有关质量和学生过度依赖LLM产生的内容存在担忧.
- 本研究评估了物理教育中LLM生成的概念项目.
研究的目的:
- 评估ChatGPT生成的概念物理项目的质量和特征.
- 将人工智能生成的项目与已建立的概念库存进行比较.
- 了解使用LLM在评估创建中对教育工作者的影响.
主要方法:
- 优化提示生成30个概念项目在动力学使用ChatGPT.
- 专家审查和选择前15项.
- 与"部队概念库存" (FCI) 一起向172名大学生提供物品.
- 对学生的答案进行了确认因素分析.
主要成果:
- 聊天GPT生成的项目显示中等难度和歧视.
- 与FCI相比,人工智能产生的项目平均表现略低.
- 确认因素分析支持与专家预期一致的三因素模型.
结论:
- 高质量的概念项目可以由LLM生成,需要大量的快速工程和选择努力.
- 人工智能产生的物品接近人类创造的物品的质量,但需要仔细审查.
- 人类监督和学生反对于完善人工智能产生的评估至关重要,特别是对于分散注意力的评估.
更多相关视频
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Self-Report Tests of Personality
310
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
310
Accuracy and Errors in Hypothesis Testing
176
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
176
Measures of Intelligence
6.6K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
6.6K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Quality Assurance
115
Quality assurance is the overarching term used to describe the activities employed to ensure the proper performance of a system. These activities can be classified into three categories: quality control, quality assessment, and internal corrective measures. Typically, these activities work cyclically: quality control is performed before and during the analysis, while quality assessment occurs during and after the investigation. Internal corrective measures are implemented based on the findings...
115


