不同物品运行对计算机适应性测试在不同条件下的影响
Merve Sahin Kursad1, Seher Yalcin2
1TED University, Turkey.
Applied psychological measurement
|October 28, 2024
概括
差异性项目功能 (DIF) 对测量精度,测试信息功能 (TIF) 和计算机自适应测试 (CAT) 的有效性产生负面影响. DIF项的百分比和类型显著降低了测量精度.
科学领域:
- 心理测量 心理测量 心理测量
- 教育测量教育的测量
- 计算机化的适应性测试
背景情况:
- 差异物品功能 (DIF) 可以危及测试成绩的有效性.
- 计算机适应性测试 (CAT) 由于其动态的项目选择,容易受到DIF的影响.
- 了解DIF的影响对于保持测试精度和有效性至关重要.
研究的目的:
- 研究DIF对测量精度,测试信息功能 (TIF) 和CAT的整体测试有效性的影响.
- 分析不同的DIF参数如何影响CAT框架内的关键心理特征.
- 量化DIF对能力估计准确性和可靠性的影响.
主要方法:
- 使用Rstudio模拟数据生成,操纵项目池大小,DIF类型和DIF百分比.
- CAT模拟包括各种项目选择和测试结束规则.
- 在不同的DIF条件下计算测量精度,TIF和测试有效性.
主要成果:
- 发现差异物品功能 (DIF) 对测量精度,TIF和测试有效性产生负面影响.
- DIF项目的百分比和DIF的类型表明对测量精度有统计学上显著的负面影响.
- 具有较高DIF项目百分比的CAT在能力估计中的准确性降低.
结论:
- DIF对CAT的心理测量完整性构成重大威胁.
- 测试开发人员必须仔细考虑和减轻DIF,以确保准确可靠的测量.
- 在CAT中对DIF检测和控制方法的进一步研究是有必要的.
更多相关视频
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
709
10:26Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
3.9K
相关概念视频
Friedman Two-way Analysis of Variance by Ranks
150
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
150
Comparing Experimental Results: Student's t-Test
1.5K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
1.5K
The Anderson-Darling Test
675
The Anderson-Darling test is a statistical method used to determine whether a data sample is likely drawn from a specific theoretical distribution. Unlike parametric tests, it does not require assumptions about specific parameters of the distribution. Instead, it compares the sample's empirical cumulative distribution function (ECDF) with the cumulative distribution function (CDF) of the hypothesized distribution. Critical values for the test are specific to the chosen distribution rather...
675
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
1.6K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.6K
Cochran's Q Test
230
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
230
