A comparison of the predictive performance of continuous and class-based latent trait models
Wanjing Anya Ma1, Yiqing Liu1, Klint Kanopka2
1Graduate School of Education, Stanford University.
None:
The ability of a student can be conceptualized as either a continuously varying entity (e.g., conventional analysis using dichotomous item response theory [IRT] models; Lord & Novick, 1968) or a bundle of latent classes (e.g., cognitive diagnostic models [CDMs]; Rupp et al., 2010; von Davier & Lee, 2019). This article builds on recent efforts to focus on predictive differences between measurement models in an attempt to examine the degree to which such approaches, which utilize quite distinctive notions regarding the nature of ability, produce different predictions of response behavior in the real world. We first present two simulation studies in which data are generated from variants of CDMs that differ in sample size, attribute hierarchical structures, and attribute estimation methods. We illustrate that, as we would expect given that they were used to generate data, CDM-based predictions uniformly outperform those of IRT models. We then compare the performance of CDM- and IRT-based approaches across 11 empirical data sets previously analyzed using CDMs. Our findings indicate that overfitting is a pervasive issue across CDM-based predictions, particularly with the generalized deterministic inputs, noisy "and" gate model. Furthermore, in no case does the CDM show superior performance when using the maximum a posteriori estimator, and only six out of 11 data sets show improved model fit for CDMs over the two-parameter logistic model when using the marginal mastery probabilities estimator. Researchers and practitioners may need to balance the diagnostic appeal of CDMs with the fact that their complexity can come at the cost of predictive accuracy. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Comparing the Survival Analysis of Two or More Groups
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Multiple Allele Traits
