Score-Based Tests of Differential Item Functioning via Pairwise Maximum Likelihood Estimation
Ting Wang1, Carolin Strobl2, Achim Zeileis3
1Department of Psychological Sciences, University of Missouri, Columbia, MO, USA. twb8d@mail.missouri.edu.
Psychometrika
|November 19, 2017
Summary
This study introduces novel score-based tests to detect violations of measurement invariance in item response theory models. These tests effectively identify problematic parameters without needing prior information, improving scale interpretation.
Area of Science:
- Psychometrics
- Statistical modeling
- Educational measurement
Background:
- Measurement invariance is crucial in item response theory (IRT) for accurate latent construct assessment.
- Violations can lead to misinterpretations and systematic bias, particularly in diverse populations.
- Existing detection methods often require unavailable prior information on item parameters and groups.
Purpose of the Study:
- To extend recently developed score-based tests for detecting measurement invariance violations.
- To adapt these tests for two-parameter item response models, focusing on pairwise maximum likelihood.
- To evaluate the tests' efficacy in identifying problematic item parameters.
Main Methods:
- Utilizing score-based tests derived from casewise derivatives of the likelihood function.
- Estimating only the null model (assuming measurement invariance holds).
- Applying tests to two-parameter IRT models with pairwise maximum likelihood estimation.
Main Results:
- The proposed score-based tests demonstrate effectiveness in identifying problematic item parameters in simulations.
- The study details the theoretical underpinnings and practical implementation of these novel tests.
- An empirical example showcases the real-world application of the measurement invariance tests.
Conclusions:
- Score-based tests offer a practical alternative for detecting measurement invariance violations in IRT.
- These tests are valuable for ensuring scale validity and fairness across different groups.
- The extension to two-parameter models broadens the applicability of these statistical tools.
Related Concept Videos
Friedman Two-way Analysis of Variance by Ranks
515
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
515
Bonferroni Test
3.4K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
3.4K
Multiple Comparison Tests
4.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.5K
Fisher's Exact Test
1.3K
Fisher's exact test is a statistical significance test widely used to analyze 2x2 contingency tables, particularly in situations where sample sizes are small. Unlike the chi-squared test, which approximates P-values and assumes minimum expected frequencies of at least five in each cell, Fisher's exact test calculates the exact probability (P-value) of observing the data or more extreme results under the null hypothesis. This feature makes it especially valuable when the assumptions of...
1.3K
Expected Frequencies in Goodness-of-Fit Tests
8.7K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.7K
Self-Report Tests of Personality
891
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
891


