The Impact and Detection of Uniform Differential Item Functioning for Continuous Item Response Models
1Ball State University, Muncie, IN, USA.
Educational and Psychological Measurement
|September 4, 2023
Summary
This study investigated methods for detecting differential item functioning (DIF) in continuous response models (CRMs). The MIMIC model demonstrated superior performance in controlling Type I errors and detecting DIF in CRMs.
Area of Science:
- Psychometrics
- Educational Measurement
- Statistical Modeling
Background:
- Item response theory (IRT) is widely used for categorical data.
- Continuous response models (CRMs) are increasingly used with computer-based testing.
- Investigating differential item functioning (DIF) in CRMs requires further research.
Purpose of the Study:
- To compare the effectiveness of three methods for assessing DIF in CRMs.
- To evaluate regression, MIMIC model, and factor invariance testing for DIF detection.
- To identify optimal methods for DIF analysis in continuous response data.
Main Methods:
- A simulation study was conducted to compare DIF detection methods.
- Methods assessed include regression, the MIMIC model, and factor invariance testing.
- Performance was evaluated based on Type I error control and statistical power.
Main Results:
- The MIMIC model exhibited effective Type I error control.
- The MIMIC model demonstrated relatively high power for detecting DIF.
- Factor invariance testing and regression showed limitations in performance.
Conclusions:
- The MIMIC model is a promising approach for DIF detection in CRMs.
- Findings inform the selection of appropriate methods for DIF analysis in continuous data.
- Further research can refine DIF detection strategies for advanced measurement models.
Related Concept Videos
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Response Surface Methodology
176
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
176
Friedman Two-way Analysis of Variance by Ranks
233
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
233
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
1.6K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.6K
Goodness-of-Fit Test
3.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.4K
Difference from Background: Limit of Detection
6.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.4K


