Assessing Measurement Invariance in ASWB Exams: Regulatory Research Proposal to Advance Equity
Matthew DeCarlo1, Gerald Bean2
1College of Education and Human Development, Saint Joseph's University, Philadelphia, Pennsylvania, USA.
Journal of Evidence-Based Social Work (2019)
|February 12, 2024
Summary
This study examined measurement invariance in social work licensing exams using simulated data. Small amounts of differential item functioning contributed to differential test functioning, indicating potential bias in exams.
Area of Science:
- Psychometrics
- Social Work Education
- Measurement Invariance
Background:
- Social workers from minoritized groups face disparities in passing licensing exams.
- Licensing exams are crucial for professional practice.
- Understanding exam fairness is essential for equitable access to the profession.
Purpose of the Study:
- To investigate the measurement equivalence (invariance) of social work licensing examinations.
- To simulate response data and analyze the impact of differential item functioning (DIF) on overall test performance.
- To assess the relationship between item-level and test-level functioning in the context of licensing exams.
Main Methods:
- Simulated response data for 15 multiple-choice questions.
- Utilized the R mirt package to fit a 2-parameter logistic (2PL) model.
- Generated data with five items exhibiting DIF and analyzed their impact on test and item characteristic curves.
Main Results:
- Small amounts of differential item functioning aggregated into differential test functioning.
- The overall effect size of DIF on test functioning was found to be small.
- This outcome represents a potential finding from analyses of social work licensing exams like the ASWB exams.
Conclusions:
- Differential test functioning is a key component of measurement invariance studies.
- Psychometric standards mandate item-level and test-level measurement invariance assessments to mitigate bias.
- Analyzing test characteristic curves aids in identifying content-irrelevant variance, such as unfairness or unreliable pass scores.
Related Concept Videos
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K
One-Way ANOVA: Unequal Sample Sizes
5.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.8K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Measures of Intelligence
7.2K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.2K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Friedman Two-way Analysis of Variance by Ranks
196
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
196


