提高预测差异物品功能的精度:一种M-DIF预训练模型方法
1Nagoya University, Japan.
Educational and psychological measurement
|November 18, 2024
概括
新的统一M-DIF模型在各种测试条件下一致量化了差异物品功能 (DIF) 大小. 这种强大的模型提高了DIF检测的准确性,为教育和心理评估提供了更可靠的方法.
科学领域:
- 心理测量 心理测量 心理测量
- 教育测量教育的测量
- 统计建模 统计建模
背景情况:
- 现有的差异物品功能 (DIF) 检测方法缺乏一致的效果大小定义和对测试条件的充分考虑.
- 不一致的DIF大小估计妨碍了对大规模评估的准确解释和应用.
研究的目的:
- 引入一个统一的M-DIF模型,以便对DIF大小进行一致和定量定义.
- 开发一个强大的模型,包括各种DIF检测方法和测试条件.
- 提高DIF分析在不同数据集和测试场景中的通用性和适用性.
主要方法:
- 开发了统一的M-DIF模型,将DIF大小定义为参考和焦点组之间的项目难度参数差异.
- 采用预训练方法,使用大型训练数据集 (测试条件的144个组合,144,000个项目,29个指标) 与XGBoost建模.
- 在一致和不一致的测试条件下,使用根平均平方误差 (RMSE) 和BIAS指标与基线模型对模型的性能进行验证.
主要成果:
- 在两个验证组中,M-DIF模型显著超过了基线模型,在一致和不一致的测试条件下显示出更高的准确性.
- 在360个测试条件组合中,M-DIF模型在99.2%的案例中表现出较低的RMSE,突出了其强度和可靠性.
- 一个实证示例证实了在现实世界评估场景中实施M-DIF模型的实际可行性和有效性.
结论:
- 统一的M-DIF模型为量化DIF大小提供了一个标准化和准确的方法,解决了现有方法的局限性.
- 该模型在各种测试条件中的稳定性使其成为提高教育和心理评估质量和公平性的宝贵工具.
- 预训练方法确保了模型的通用性,允许直接应用于新数据,并促进更可靠的DIF检测.
相关概念视频
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Identifying Statistically Significant Differences: The F-Test
1.6K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
1.6K


