对WAIC进行分解,以评估信息获取,并应用于教育测试
Fang Liu1, Ming-Hui Chen2, Xiaojing Wang2
1Soochow University, Suzhou, Jiangsu, China.
The British journal of mathematical and statistical psychology
|February 21, 2025
概括
这项研究引入了一种新的方法来评估额外的教育测试数据,如响应时间,是否可以改善项目响应模型. 这些发现有助于确定适合教育测试数据的最有用的维度.
科学领域:
- 教育测量教育的测量
- 心理测量 心理测量 心理测量
- 统计建模 统计建模
背景情况:
- 多维数据在教育测试中越来越常见.
- 评估额外数据维度对安装项目响应数据的有用性至关重要.
- 现有的模型评估标准可能无法充分利用多维信息.
研究的目的:
- 使用后预测坐标 (PPO) 开发一个新的广泛适用信息标准 (WAIC) 的分解.
- 提出一种新的模型评估标准,用于评估项目响应模型中多维数据的有用性.
- 确定响应时间和其他教育成绩在匹配响应数据中的相对重要性.
主要方法:
- 开发了一个联合模型,包括响应,响应时间和额外的教育分数.
- 提出了一个新的模型评估标准,该标准基于通过PPO进行WAIC分解.
- 采用高效的蒙特卡洛方法计算PPO.
- 进行了广泛的模拟,并分析了来自计算机评估程序的真实数据集.
主要成果:
- 拟议的模型评估标准有效地确定了适配响应数据最有用的维度.
- 证明了额外的维度,如响应时间,可以显著提高模型的适应性.
- 蒙特卡洛方法提供了一种有效的方式来计算复杂模型的PPO.
结论:
- 新的标准有助于选择最佳的多维模型进行教育测试.
- 响应时间和其他教育分数为增强项目响应模型提供了宝贵的信息.
- 该方法是稳固的,适用于现实世界的教育评估数据.
相关概念视频
Wechsler's Contribution to Measures of Intelligence
1.4K
David Wechsler, a psychologist who worked with World War I veterans, developed a significant IQ test in 1939 called the Wechsler-Bellevue Intelligence Scale. This test was innovative because it combined several subtests that measured both verbal and nonverbal skills, reflecting Wechsler's belief that intelligence is a global capacity involving purposeful action, rational thinking, and effective interaction with the environment. This test later evolved into the Wechsler Adult Intelligence...
1.4K
Wald-Wolfowitz Runs Test I
592
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
592
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Kendall's Coefficient of Concordance
215
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
215
Wilcoxon Rank-Sum Test
137
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
137
Review and Preview
6.9K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
6.9K


