Related Experiment Video
Updated: Jun 23, 2025

08:27
Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
Published on: September 27, 2019
6.9K
The Accuracy of Bayesian Model Fit Indices in Selecting Among Multidimensional Item Response Theory Models.
1Loyola University Chicago, IL, USA.
Educational and Psychological Measurement
|June 20, 2024
Summary
Bayesian model comparison for Item Response Theory (IRT) data can favor complex nested models. Using appropriate indices like Pareto-smoothed importance sampling avoids bias, ensuring accurate dimensionality assessment.
Area of Science:
- Psychometrics
- Statistical Modeling
- Educational Measurement
Background:
- Item Response Theory (IRT) models are crucial for analyzing rating scale data and determining dimensionality.
- Model comparisons based on predictive performance can be biased, favoring nested-dimensionality IRT models over non-nested ones.
- The extent of this bias, particularly with Bayesian estimation and specific comparison indices, remains unclear.
Purpose of the Study:
- To investigate the bias in Bayesian predictive performance indices when comparing nested- and non-nested-dimensionality IRT models.
- To assess the accuracy of four Bayesian indices in differentiating between these model types under various dimensional structures.
- To clarify the conditions under which nested-dimensionality models might be unfairly favored.
Main Methods:
- A simulation study was conducted to evaluate Bayesian predictive performance indices.
- Four indices were examined: Deviance Information Criterion (DIC), Pareto-smoothed importance sampling (PSIS-LOO), Watanabe-Akaike Information Criterion (WAIC), and log-predicted marginal likelihood (LPML).
- The study simulated data representing specific dimensional structures to test model differentiation accuracy.
Main Results:
- The Deviance Information Criterion (DIC) showed extreme bias, favoring nested-dimensionality models even when incorrect.
- Pareto-smoothed importance sampling (PSIS-LOO) demonstrated the least bias among the tested indices.
- Watanabe-Akaike Information Criterion (WAIC) and log-predicted marginal likelihood (LPML) also performed relatively well, with less bias than DIC.
Conclusions:
- Nested-dimensionality IRT models are not inherently favored when data represent specific dimensional structures if appropriate Bayesian predictive indices are used.
- The choice of model comparison index is critical for accurate dimensionality assessment in IRT.
- Pareto-smoothed importance sampling (PSIS-LOO) is recommended for reliable model comparison in these contexts.
Related Concept Videos
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Friedman Two-way Analysis of Variance by Ranks
180
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
180
Response Surface Methodology
117
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
117
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
1.6K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.6K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K

