一种多维贝叶斯式IRT方法,用于从概念测试数据中发现误解
Martin Segado1, Aaron Adair2, John Stewart3
1Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA, United States.
Frontiers in psychology
|February 21, 2025
概括
这项研究引入了一种新方法,可以自动识别学生在多项选择测试中的误解. 该方法成功地发现了物理学中已知的和潜在的新学习困难,有助于有针对性的教学.
科学领域:
- 教育测量教育的测量
- 认知科学 认知科学
- 物理教育研究 物理学教育研究
背景情况:
- 识别学生的误解对于有效的教学至关重要.
- 传统方法通常需要对测试数据进行手动分析.
- 现有的心理测量模型可能无法完全捕捉学生思维的细微差别.
研究的目的:
- 开发和验证一种探索性方法,用于从多选项概念测试中发现学生的误解.
- 为了利用贝叶斯实现的多维名义类别项目响应理论模型 (MNCM) 来自动识别误解.
- 证明该方法在没有手动内容标签的情况下识别误解的能力.
主要方法:
- 利用贝叶斯的MNCM与因子分析旋转方法相结合.
- 在个别分心因素水平上分析了学生的反应.
- 在合成数据上验证了该方法,并将其与现有的IRT软件进行了比较.
主要成果:
- 贝叶斯式MNCM准确地从合成数据中恢复了多维项目参数.
- 该方法显示了对过度装配的稳定性,并进行了自动维度评估.
- 应用于"力量概念库存",该方法确定了总得分以外的13个额外维度,与已知的牛顿力学误解一致.
结论:
- 开发的方法有效地从概念测试数据中发现可能存在的学生误解.
- 这种方法可以完善现有的概念库存或帮助开发新的概念库存.
- 调查结果支持发现新的误解和实现有针对性的指导的潜力.
相关概念视频
Multiple Comparison Tests
3.8K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.8K
Accuracy and Errors in Hypothesis Testing
169
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
169
Introduction to Test of Independence
2.1K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.1K
Cause and Effect
10.8K
While variables are sometimes correlated because one does cause the other, it could also be that some other factor, a confounding variable, is actually causing the systematic movement in our variables of interest. For instance, as sales in ice cream increase, so does the overall rate of crime. Is it possible that indulging in your favorite flavor of ice cream could send you on a crime spree? Or, after committing crime do you think you might decide to treat yourself to a cone?
10.8K
Hypothesis Test for Test of Independence
3.4K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.4K
Errors In Hypothesis Tests
4.1K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
4.1K


