Approximate Functional Relationship between IRT and CTT Item Discrimination Indices: A Simulation, Validation, and
John T Kulas1, Jeffrey A Smith, Hui Xu
1John T. Kulas, 1046 Voyageur St., St. Cloud, MN 56303, USA, jtkulas@stcloudstate.edu.
Summary
This study refines Lord's equation to link classical test theory (CTT) and item response theory (IRT) discrimination indices. The modified formula, incorporating item difficulty, offers improved practical application and prediction accuracy for educational and workforce testing.
Area of Science:
- Psychometrics
- Educational Measurement
- Psychological Statistics
Background:
- Classical Test Theory (CTT) and Item Response Theory (IRT) are fundamental measurement paradigms.
- Lord (1980) proposed a conceptual equation linking CTT and IRT discrimination indices.
- Existing methods lack practical utility and empirical validation for this linkage.
Purpose of the Study:
- To modify Lord's (1980) conceptual equation for practical application.
- To develop an empirical function relating CTT and IRT discrimination indices.
- To incorporate item difficulty and corrected item-total correlation into the linkage.
Main Methods:
- Simulation of over 768 trillion item responses to determine a best-fitting empirical function.
- Modification of Lord's (1980) equation to include item difficulty.
- Validation using real-world data from 16 diverse workforce and educational tests.
Main Results:
- The modified equation demonstrates shifted functional asymptotes, slopes, and points of inflection.
- The empirical function provides good prediction accuracy under standard psychometric assumptions (e.g., normal ability distribution, moderate difficulty).
- The proposed modification offers greater precision compared to Lord's (1980) original formula.
Conclusions:
- The refined equation provides a practical and empirically validated method for relating CTT and IRT discrimination indices.
- This advancement enhances the practical utility of psychometric linkage for test developers and researchers.
- The findings support the use of the modified function in educational and workforce assessment contexts.
Related Concept Videos
Multiple Comparison Tests
4.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.5K
Regression Toward the Mean
7.2K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.2K
Kendall's Tau Test
1.2K
Kendall's tau test, also known as the Kendall rank coefficient test, is a nonparametric method for assessing association between two variables. This test is particularly useful for identifying significant correlations when the distributions of the sample and population are unknown. Developed in 1938 by the British statistician Sir Maurice George Kendall, the tau coefficient (denoted as τ) serves as a rank correlation coefficient, with values ranging from -1 to +1.
A τ value of +1 indicates...
A τ value of +1 indicates...
1.2K
Comparing Experimental Results: Student's t-Test
6.1K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
6.1K
Expected Frequencies in Goodness-of-Fit Tests
8.7K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.7K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
7.0K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
7.0K


