Fit Indices for Measurement Invariance Tests in the Thurstonian IRT Model
1California State University Channel Islands, Camarillo, USA.
Applied Psychological Measurement
|June 16, 2020
Summary
This study found that traditional fit index cutoffs poorly detect measurement invariance in Thurstonian item response theory (IRT) models. New, stricter cutoffs for ΔCFI and ΔNCI are recommended for accurate invariance testing in forced-choice formats.
Area of Science:
- Psychometrics
- Statistical Modeling
Background:
- Traditional fit indices for confirmatory factor analysis (CFA) are widely used but their applicability to Thurstonian item response theory (IRT) models, especially for multiple-group analyses, remains unclear.
- Assessing measurement invariance is crucial for cross-cultural comparisons and ensuring test fairness, particularly with forced-choice formats.
Purpose of the Study:
- To evaluate the effectiveness of existing fit index cutoffs for assessing model fit and testing measurement invariance in Thurstonian IRT models.
- To investigate the impact of measurement non-invariance on score estimation accuracy and efficiency within the Thurstonian IRT framework.
- To propose refined cutoffs for fit indices to improve the detection of metric and scalar non-invariance in forced-choice data.
Main Methods:
- Multiple group confirmatory factor analysis (CFA) was employed within the Thurstonian IRT model framework.
- The performance of various fit index cutoffs was examined by analyzing detection rates of measurement non-invariance and Type I error rates.
- Sampling distributions of fit index differences were generated to derive new suggested cutoffs.
Main Results:
- Existing fit index cutoffs showed limited success in detecting metric non-invariance and failed to detect scalar non-invariance.
- New cutoffs, specifically ΔCFI > .001 and ΔNCI > .004 for scalar non-invariance, and ΔCFI > .007 for metric non-invariance, demonstrated better performance.
- ΔCFI was recommended as the most suitable index for measurement non-invariance tests in forced-choice data, balancing error control and detection rates.
Conclusions:
- Traditional fit index cutoffs are inadequate for robust measurement invariance testing in Thurstonian IRT models.
- Revised fit index criteria are necessary to accurately assess invariance in forced-choice formats, enhancing cross-cultural test development.
- Further research is needed to refine these methods and improve the utility of forced-choice formats in international testing contexts.
Related Concept Videos
Inertia Tensor
959
The concept of the inertia tensor is employed to depict the mass distribution and rotational inertia of a solid or rigid object. This tensor is expressed through a three-by-three matrix. Each component within this matrix corresponds to varying moments of inertia about specific axes.
The diagonal components of the inertia tensor matrix represent the moments of inertia concerning the principal axes of the object. These primary axes are defined as the axes where the object experiences the least...
The diagonal components of the inertia tensor matrix represent the moments of inertia concerning the principal axes of the object. These primary axes are defined as the axes where the object experiences the least...
959
One-Way ANOVA: Unequal Sample Sizes
6.5K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
6.5K
Friedman Two-way Analysis of Variance by Ranks
426
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
426
Expected Frequencies in Goodness-of-Fit Tests
6.3K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
6.3K
Measures of Intelligence
8.1K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.1K
Comparing Experimental Results: Student's t-Test
4.5K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.5K


