Related Experiment Video
Updated: Aug 29, 2026

Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
A Comparison of Item Selection Methods and Parameter Estimation Approaches for Online Calibration in Computerized
Levent Ertuna1, Cees A W Glas2, Burcu Atar3
1Sakarya University, Türkiye.
Abstract:
Online calibration enables continuous replenishment of computerized adaptive testing (CAT) item banks by embedding new pretest items within operational test sessions. However, the relative performance of different calibration components and their interactions have not been fully examined in a unified framework. This study compared three pretest item selection methods (maximum Fisher information [MFI], D-optimal value design, and Bayesian D-optimal design), two parameter estimation methods (a fixed-ability joint maximum likelihood procedure [JML-itpar] and one expectation-maximization cycle [OEM]), three levels of random calibration stage sample size (250, 500, and 1,000), and three levels of calibration sample size per pretest item (250, 500, and 1,000) under the one-parameter logistic and two-parameter logistic item response theory models through a Monte Carlo simulation with 108 conditions and 100 replications per condition. The results revealed a trade-off among item selection methods: MFI yielded the most accurate difficulty parameter estimates, while D-optimal value design and Bayesian D-optimal design produced better discrimination parameter recovery. Bayesian D-optimal design offered the best balance between accuracy and calibration efficiency. For the first time in the standard online calibration context, JML-itpar was systematically evaluated and found to be a viable estimation method. JML-itpar paired with MFI produced lower discrimination parameter root mean squared error than OEM at calibration sample sizes of 500 and 1,000. OEM remained superior for the difficulty parameter. The calibration sample size per pretest item was the most influential factor, with substantial accuracy gains from 250 to 500 responses and diminishing returns beyond 500. The random calibration stage sample size had negligible effects. These findings provide practical guidance for selecting calibration components in operational CAT programs.
Related Concept Videos
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with data...
Methods of Medium Optimization
Instrument Calibration
Analytical Balance Calibration
An analytical balance measures mass and requires regular calibration to...

