Evaluation of MIMIC-Model Methods for DIF Testing With Comparison to Two-Group Analysis.
1a Washington University in St. Louis .
Multivariate Behavioral Research
|January 23, 2016
Summary
Differential item functioning (DIF) detection is crucial for test fairness. Multiple-indicator multiple-cause (MIMIC) models offer a more accurate approach than traditional 2-group item response theory (IRT) methods, especially with small focal groups.
Area of Science:
- Psychometrics
- Educational Measurement
- Statistical Modeling
Background:
- Differential item functioning (DIF) is critical for ensuring test fairness across diverse groups.
- Existing DIF detection methods, like 2-group item response theory (IRT), have limitations, particularly with small focal groups.
- The accuracy and sample size requirements for multiple-indicator multiple-cause (MIMIC) models in DIF testing are not well-established.
Purpose of the Study:
- To examine the accuracy of MIMIC structural equation models for DIF testing when the focal group sample size is small.
- To compare the performance of MIMIC models against traditional 2-group IRT methods for DIF detection.
- To provide recommendations for applying MIMIC methods in DIF testing scenarios with limited focal group data.
Main Methods:
- Utilized multiple-indicator multiple-cause (MIMIC) structural equation models, parameterized as item response models.
- Employed simulation studies to assess the accuracy of MIMIC methods under small focal group conditions.
- Compared MIMIC model results with those from 2-group item response theory (IRT) analyses.
Main Results:
- MIMIC models demonstrated superior accuracy in detecting uniform DIF compared to 2-group IRT, particularly with small focal-group samples.
- The accuracy of MIMIC methods was validated for both binary and 5-category ordinal response data.
- Results support the utility and robustness of the MIMIC approach for DIF testing in challenging sample size situations.
Conclusions:
- The MIMIC modeling approach is a valuable and accurate tool for differential item functioning (DIF) testing, especially when dealing with small focal groups.
- Findings suggest that MIMIC models provide more reliable DIF detection than 2-group IRT under these conditions.
- The study offers practical guidance for researchers and test developers on implementing MIMIC methods for enhanced measurement fairness.
Related Concept Videos
Multiple Comparison Tests
4.6K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.6K
Comparing Experimental Results: Student's t-Test
6.4K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
6.4K
Friedman Two-way Analysis of Variance by Ranks
561
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
561
Test for Homogeneity
2.5K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.5K
Identifying Statistically Significant Differences: The F-Test
4.2K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
4.2K
Bonferroni Test
3.5K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
3.5K


