Related Experiment Video
Updated: Jun 28, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Combining regularization and logistic regression model to validate the Q-matrix for cognitive diagnosis model
Xiaojian Sun1,2,3, Tongxin Zhang4, Chang Nie4
1School of Mathematics and Statistics, Southwest University, Chongqing, China.
Abstract:
Q-matrix is an important component of most cognitive diagnosis models (CDMs); however, it mainly relies on subject matter experts' judgements in empirical studies, which introduces the possibility of misspecified q-entries. To address this, statistical Q-matrix validation methods have been proposed to aid experts' judgement. A few of these methods, including the multiple logistic regression-based (MLR-B) method and the Hull method, can be applied to general CDMs, but they are either time-consuming or lack accuracy under certain conditions. In this study, we combine the L1 regularization and MLR model to validate the Q-matrix. Specifically, an L1 penalty term is imposed on the log-likelihood of the MLR model to select the necessary attributes for each item. A simulation study with various factors was conducted to examine the performance of the new method against the two existing methods. The results show that the regularized MLR-B method (a) produces the highest Q-matrix recovery rate (QRR) and true positive rate (TPR) for most conditions, especially with a small sample size; (b) yields a slightly higher true negative rate (TNR) than either the MLR-B or the Hull method for most conditions; and (c) requires less computation time than the MLR-B method and similar computation time as the Hull method. A real data set is analysed for illustration purposes.
More Related Videos
Related Concept Videos
Detection of Gross Error: The Q Test
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Regression Toward the Mean
Cochran's Q Test
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

