Related Experiment Video
Updated: Aug 6, 2026

Problem-Solving Before Instruction (PS-I): A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
Consensus Among Differential Item Functioning Effect Size Measures: A Simplified Approach to Reporting Effect Size
1Department of Psychology, University of Notre Dame, IN, USA.
Abstract:
Differential item functioning (DIF) is an issue of a measure that can lead to differences in item scores between subpopulations (e.g., sex, race, age) while controlling for a latent trait or ability level. It is important to report DIF effect size measures to understand the practical significance of DIF results as well. We reviewed 15 different DIF effect size measures in the literature, and no previous study had examined whether the 15 DIF effect size measures agreed with one another, which has implications in the number and type of measures to report. The present study investigated the relationship among DIF effect size measures in 1PL, 2PL, and GRM models and provide clear recommendations to applied researchers for their reporting. Our simulation results showed that, for each model, three effect size measures can be reported to obtain the most unique information about uniform DIF and the least bias from sampling conditions (e.g., sample size and impact): Mantel-Haenszel Delta, McFadden's Pseudo R-squared, and a signed DIF measure. However, the optimal signed DIF measure varies by model, with Signed Area under the curve for the 1PL model, Standardized P-difference for the 2PL, and Signed Item Difference in the Sample/Normal Distribution for the GRM. The recommended approach to reporting DIF effect size improves the detection of DIF through practical significance and provides researchers with more tools to address bias.
Related Concept Videos
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Friedman Two-way Analysis of Variance by Ranks
Odds Ratio
One-Way ANOVA: Unequal Sample Sizes
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Behrens–Fisher Test
This test is...

