调整后的剩余值用于评估多阶段适应性测试的IRT模型中的条件独立性
Peter W van Rijn1, Usama S Ali2,3, Hyo Jeong Shin4
1ETS Global.
Psychometrika
|February 25, 2026
概括
项目响应理论 (IRT) 模型需要根据路由决策对多阶段适应性测试 (MST) 数据进行调整. 未调整的余量是不合适的,但调整的余量在模拟和真实数据分析中显示出令人满意的I型错误.
科学领域:
- 心理测量 心理测量 心理测量
- 统计建模 统计建模
- 教育测量教育的测量
背景情况:
- 项目响应理论 (IRT) 模型假设项目响应的条件独立性给定潜能.
- 多阶段适应性测试 (MST) 设计涉及路由决策,可能会违反这种有条件独立性假设.
- 这种违规行为可能会影响心理测量分析中的统计推断.
研究的目的:
- 研究MST设计中的路由决定对IRT模型中条件独立性假设的影响.
- 评估一般化残留物对于分析MST数据的适当性.
- 建议和验证对MST数据的残留量进行调整.
主要方法:
- 该研究在MST数据的背景下研究了准独立现象.
- 它证明了标准通用余量 (Haberman & Sinharay,2013) 对未经修改的MST数据的不适用性.
- 根据MST设计的复杂性,对残留物进行调整,并通过模拟和真实数据分析进行评估.
主要成果:
- 对于项目对频率的概括余数在没有设计特定调整的情况下,对于MST数据是不合适的.
- 拟议的调整后的残余值在模拟研究中显示出令人满意的I型错误率.
- 调整后的剩余值成功地应用于国际学生评估计划 (PISA) 的真实MST数据.
结论:
- 标准的IRT残留诊断不能直接适用于MST数据.
- 对残余的调整对于MST设计中有效的统计推断是必要和有效的.
- 这些发现对适应性测试环境中的统计推断和模型合适性评估有影响.
更多相关视频
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.9K
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
1.3K
相关概念视频
Introduction to Test of Independence
3.0K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
3.0K
Friedman Two-way Analysis of Variance by Ranks
522
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
522
Hypothesis Test for Test of Independence
8.3K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
8.3K
Assumptions of Survival Analysis
462
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
462
Determination of Expected Frequency
2.6K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.6K
Expected Frequencies in Goodness-of-Fit Tests
8.7K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.7K
