Improving the assessment of measurement invariance: Using regularization to select anchor items and identify
William C M Belzak1, Daniel J Bauer1
1Department of Psychology and Neuroscience, University of North Carolina at Chapel Hill.
Psychological Methods
|January 10, 2020
Summary
Lasso regularization effectively identifies differential item functioning (DIF) and selects anchor items in measurement invariance testing. This machine learning method offers superior control over Type I errors compared to traditional likelihood ratio tests, especially with large sample sizes and substantial DIF.
Area of Science:
- Behavioral Sciences
- Psychometrics
- Statistics
Background:
- Evaluating measurement invariance is crucial in behavioral sciences.
- Measurement invariance fails when differential item functioning (DIF) exists, where item responses differ across groups.
- Traditional DIF detection methods iteratively test items, risking incorrect anchor selection and DIF identification.
Purpose of the Study:
- To propose and evaluate an alternative method for selecting anchors and identifying DIF using regularization.
- Specifically, to apply a lasso penalty within the two-parameter logistic item response theory model.
- To compare the performance of lasso regularization against the likelihood ratio test method for DIF analysis.
Main Methods:
- Utilized regularization, a machine learning technique, specifically a lasso penalty.
- Applied the method to group differences in item parameters within the two-parameter logistic item response theory model.
- Compared lasso regularization with the likelihood ratio test method in a 2-group DIF analysis using simulations and empirical data.
Main Results:
- Lasso regularization demonstrated superior control of Type I error compared to the likelihood ratio test method.
- This advantage was particularly evident when substantial DIF was present and sample sizes were large.
- The power of lasso regularization showed only a minor decrement, indicating robust performance.
Conclusions:
- Lasso regularization is a promising alternative for testing DIF and selecting anchor items.
- The method offers improved accuracy and reliability in measurement invariance assessment.
- It provides a more robust approach, especially in complex scenarios with widespread DIF.
Related Concept Videos
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Friedman Two-way Analysis of Variance by Ranks
457
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
457
One-Way ANOVA: Equal Sample Sizes
3.9K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.9K
One-Way ANOVA: Unequal Sample Sizes
6.5K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
6.5K
The Anchoring-and-Adjustment Heuristic
7.7K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.7K
Measures of Intelligence
8.2K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.2K


