项目评估与EM算法的输出相匹配:基于后期预期的RMSD指数
Yun-Kyung Kim1, Li Cai1, YoungKoung Kim2
1University of California, Los Angeles, CA, USA.
Educational and psychological measurement
|October 7, 2025
概括
这项研究改进了使用后期期望的项目响应理论中的项目合适性分析. 一种新的切断值方法比传统的模型合适性评估方法提供了更好的准确性.
科学领域:
- 心理测量 心理测量 心理测量
- 统计建模 统计建模
- 教育测量教育的测量
背景情况:
- 项目响应理论 (IRT) 依赖于项目合适分析来确定模型的有效性.
- 后期预期 (伪计数) 为项目匹配提供了优势,特别是在缺少数据的情况下.
- 根平均平方偏差 (RMSD) 指数的解释性需要改进.
研究的目的:
- 为了提高RMSD指数的可解释性,该指数来自IRT的后期期望.
- 评估和比较不同的方法来确定物品合适度的最佳切断值.
- 根据样本大小和测试长度,开发一个可概括的RMSD参考值预测模型.
主要方法:
- 利用穷人的后部预测模型检查 (PP-PPMC) 来评估显著性水平.
- 用人接收机运行特征 (ROC) 曲线分析以经验性地确定最佳的切断值.
- 应用响应表面分析来创建参考值的预测模型.
主要成果:
- 截止值方法表现出比PP-PPMC更好的表现,有效地平衡了虚假和真正阳性率.
- 针对各种样本大小和测试长度,确定了最佳参考值.
- 预测模型准确地概括了参考值如何根据数据集特征变化.
结论:
- 该研究验证了PP-PPMC用于项目合适诊断,并引入了一种实际的频率学方法来导出参考值.
- 开发的预测模型允许研究人员计算数据集特定的RMSD参考值.
- 这项工作为IRT建模中的项目合适性分析提供了更精细的方法.
相关概念视频
Expected Frequencies in Goodness-of-Fit Tests
7.2K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
7.2K
Empirical Method to Interpret Standard Deviation
9.3K
The empirical rule, also known as the three-sigma rule, allows a statistician to interpret the standard deviation in a normally distributed dataset. The rule states that 68% of the data lies within one standard deviation from the mean, 95% lies within two standard deviations from the mean, and 99.7% lies within three standard deviations from the mean. Additionally, this rule is also called the 68-95-99.7 rule.
This rule is used widely in statistics to calculate the proportion of data values...
This rule is used widely in statistics to calculate the proportion of data values...
9.3K
Response Surface Methodology
604
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
604
Mean Absolute Deviation
3.3K
The mean absolute deviation is also a measure of the variability of data in a sample. It is the absolute value of the average difference between the data values and the mean.
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
3.3K
Goodness-of-Fit Test
8.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
8.1K
Residuals and Least-Squares Property
9.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.1K


