A Regression Discontinuity Design Framework for Controlling Selection Bias in Evaluations of Differential Item
Natalie A Koziol1, J Marc Goodrich2, HyeonJin Yoon1
1University of Nebraska-Lincoln, USA.
Educational and Psychological Measurement
|November 3, 2022
Summary
This study introduces a new method for differential item functioning (DIF) analysis that reduces bias in test accommodations. While it controls errors better, it needs more power and precision for practical use.
Area of Science:
- Psychometrics
- Educational Measurement
- Statistical Modeling
Background:
- Differential item functioning (DIF) is crucial for assessing test validity, especially with alternate form accommodations.
- Traditional DIF methods often suffer from selection bias, compromising accuracy.
- Validating test accommodations requires robust DIF analysis to ensure fairness.
Purpose of the Study:
- To propose and evaluate a novel DIF framework using regression discontinuity design to mitigate selection bias.
- To compare the proposed framework against traditional logistic regression in a simulation study.
- To assess Type I error, power, bias, and precision of DIF statistics and effect size estimators.
Main Methods:
- A simulation study was conducted to compare a new DIF framework with traditional logistic regression.
- The regression discontinuity design approach was employed to control for selection bias.
- Key metrics evaluated included Type I error rates, power, bias, and root mean square error.
Main Results:
- The novel DIF framework demonstrated superior control over Type I error rates compared to traditional methods.
- The new framework exhibited minimal bias in effect size estimation.
- However, the proposed method showed lower statistical power and precision.
Conclusions:
- The regression discontinuity-based DIF framework offers a promising approach to reduce selection bias in test accommodation validity studies.
- While effective in controlling errors, further methodological refinement is needed to enhance power and precision for practical application.
- The findings have implications for improving the fairness and validity of standardized testing accommodations.
Related Concept Videos
Friedman Two-way Analysis of Variance by Ranks
274
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
274
Group Design
9.2K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
9.2K
Experimental Designs
11.7K
An experimental design is a systematic process that allows researchers to evaluate the relationship between dependent and independent variables. There are three widely used types of experimental design - pre-experimental design, true experimental design, and quasi-experimental design. In pre-experimental design, the researcher compares the data before and after some interventions or treatments. The true-experimental design has more than one purposefully created group, a commonly measured...
11.7K
Crossover Experiments
2.9K
Crossover experiments, also called the repeated-measurements design, is a study design in which all experimental units are exposed to all treatments in different periods. Crossover experiments are generally used in psychology, the pharmaceutical industry, agriculture, and medicine.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
2.9K
Factorial Design
13.1K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.1K
Study Design in Statistics
8.4K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
8.4K


