Comparing score tests and other local dependence diagnostics for the graded response model.
1Department of Psychology, The University of North Carolina, Chapel Hill, North Carolina, USA.
This study generalizes score tests for local dependence (LD) to the graded response model, finding they offer superior power for detecting specific LD patterns compared to existing methods, especially when correctly specified. While accurate with large samples, small-sample performance is less robust than alternative tests.
Area of Science:
- Psychometrics
- Statistical modeling
- Educational measurement
Background:
- Score tests are crucial for identifying local dependence (LD) in item response theory (IRT) models.
- Existing methods primarily focus on binary item response models.
- Generalizing these tests to polytomous data, like the graded response model, is essential for broader applicability.
Purpose of the Study:
- To generalize bifactor and threshold shift score tests for local dependence detection to the graded response model.
- To evaluate the performance of these generalized score tests through simulation studies.
- To compare the power and Type I error rates of the proposed tests against established methods like Pearson's Χ2 and M2.
Main Methods:
- Generalization of bifactor score test by adding a secondary dimension for item pairs.
- Development of multiple generalizations for the threshold shift score test: conditional, uniform, and linear.
- Conducting simulation studies to assess Type I error rates and statistical power under various conditions.
- Comparison with Pearson's Χ2 and M2 statistics for local dependence detection.
Main Results:
- Generalized score tests demonstrate accurate Type I error rates in large samples.
- All proposed score tests exhibit higher power in detecting LD consistent with their parametric form, outperforming Χ2 and M2.
- Even misspecified score tests show superior power compared to Χ2 and M2 in most simulated scenarios.
- Small-sample performance of the generalized score tests is less accurate than Pearson's Χ2 and M2.
Conclusions:
- The generalized score tests are powerful tools for detecting local dependence in graded response models, particularly when the LD structure aligns with the test's parametric form.
- While effective with sufficient sample sizes, caution is advised for small-sample applications due to performance limitations compared to established alternatives.
- The study provides a valuable extension of LD detection methods for more complex item response models.
More Related Videos
Related Concept Videos
Expected Frequencies in Goodness-of-Fit Tests
Goodness-of-Fit Test
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
Comparing the Survival Analysis of Two or More Groups
Friedman Two-way Analysis of Variance by Ranks
Quantifying and Rejecting Outliers: The Grubbs Test


