Score-based measurement invariance checks for Bayesian maximum-a-posteriori estimates in item response theory
Rudolf Debelak1, Samuel Pawel2, Carolin Strobl1
1Department of Psychology, University of Zurich, Switzerland.
Summary
This study introduces new score-based statistical tests for item response theory (IRT) models, evaluating their effectiveness for Bayesian and multiple-group analyses. The pooled variance method proved practical, while simulation methods require larger sample sizes.
Area of Science:
- Psychometrics
- Educational Measurement
- Statistical Modeling
Background:
- Score-based tests assess parameter invariance in Item Response Theory (IRT) models.
- Existing tests are primarily developed within a maximum likelihood (ML) framework.
- There is a need for analogous tests applicable to Bayesian (MAP) estimates and multiple-group IRT models.
Purpose of the Study:
- To propose and evaluate score-based statistical tests for Bayesian MAP estimates and multiple-group IRT models.
- To investigate the performance of these tests against differential item functioning (DIF) using categorical and continuous covariates.
- To compare two novel test families: one based on pooled variance and another on simulation.
Main Methods:
- Development of two families of statistical tests: pooled variance approximation and simulation-based approach.
- Evaluation through a simulation study examining sensitivity to DIF.
- Application to two- and three-parameter logistic IRT models.
Main Results:
- The pooled variance method demonstrated practical utility for both ML and MAP estimates.
- The simulation-based approach showed satisfactory results only with large sample sizes.
- Both methods were evaluated for their sensitivity to DIF with various person covariates.
Conclusions:
- The pooled variance method offers a practical approach for assessing parameter invariance in IRT, including Bayesian contexts.
- The simulation-based approach requires substantial sample sizes for reliable DIF detection.
- These findings contribute to robust psychometric analysis in IRT, particularly for complex models and estimation methods.
Related Concept Videos
Friedman Two-way Analysis of Variance by Ranks
310
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
310
Expected Frequencies in Goodness-of-Fit Tests
2.7K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.7K
Measures of Intelligence
7.9K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.9K
One-Way ANOVA: Equal Sample Sizes
3.5K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.5K
Reliability and Validity
13.2K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.2K
Goodness-of-Fit Test
4.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.1K


