Related Experiment Video
Updated: Mar 8, 2026

10:58
Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics BM-PROMA
Published on: August 28, 2021
5.1K
Detecting differential item functioning in behavioral indicators across parallel forms
Juana Gómez-Benito1, Nekane Balluerka, Andrés González
1Universidad de Barcelona.
Psicothema
|January 28, 2017
Summary
This study introduces a new method using Differential Item Functioning (DIF) to assess parallel test forms. The differential functioning of behavioral indicators (DFBI) offers unique insights into item parallelism, complementing traditional total score analyses.
Area of Science:
- Psychometrics
- Educational Measurement
- Psychological Assessment
Background:
- Classical Test Theory (CTT) relies on parallel test forms, but direct verification is limited by unobservable true scores.
- Traditional methods for assessing test form parallelism face inherent limitations.
- Differential Item Functioning (DIF) offers a novel framework to address these limitations.
Purpose of the Study:
- To propose and illustrate a novel approach for verifying the parallelism of test forms using DIF.
- To shift the focus from total test scores to individual item performance in assessing parallelism.
- To introduce the "differential functioning of behavioral indicators" (DFBI) as a method to analyze item-level parallelism.
Main Methods:
- Applied several DIF techniques to analyze the performance of a single group on parallel items.
- Examined 18 items from two parallel forms of the Attention Deficit-Hyperactivity Disorder Scale.
- Utilized a dataset of 527 participants to illustrate the proposed approach.
Main Results:
- 12 out of 18 items (66.6%) exhibited statistically significant differential functioning (p < .01) based on the Mantel χ² statistic.
- Standardization revealed that DIF items were equally distributed, with half favoring Form A and half favoring Form B.
- This indicates potential non-parallel functioning at the item level despite overall test form equivalence.
Conclusions:
- The "differential functioning of behavioral indicators" (DFBI) provides valuable item-level information on test form parallelism.
- DFBI complements traditional analyses based on total scores, offering a more nuanced understanding of equivalence.
- This approach enhances the psychometric rigor in evaluating parallel test forms.
Related Concept Videos
Group Design
11.0K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
11.0K
Behrens–Fisher Test
309
The Behrens-Fisher test is a statistical method designed to address the Behrens-Fisher problem, which arises when comparing the means of two normally distributed populations with unequal variances. Unlike the Student's t-test, which assumes equal variances, the Behrens-Fisher test allows for mean comparison without this restrictive assumption. This flexibility makes it particularly valuable in scenarios where two independent samples exhibit normality but lack variance homogeneity.
This test...
This test...
309
Sign Test for Matched Pairs
448
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in...
To conduct the sign test, we first calculate the differences in...
448
One-Way ANOVA: Equal Sample Sizes
4.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.3K
Friedman Two-way Analysis of Variance by Ranks
530
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
530
Factorial Design
14.9K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
14.9K

