Related Experiment Video
Updated: Jan 29, 2026

A New Toolkit for Evaluating Gene Functions using Conditional Cas9 Stabilization
Published on: September 2, 2021
The stability of IRT parameters under several test equating conditions
Dominik Weber1, Nicolas Becker1, Frank M Spinath1
1Department of Individual Differences & Psychodiagnostics, Saarland University, Saarbrücken, Germany.
Introduction:
It is crucial for researchers and test developers to compare results from different test sets (e. g., re-testing, parallel test forms). To ensure comparability, test sets are often linked using anchor items as a common denominator alongside distinct items. To date, most studies on test equating have been limited in scope, typically comparing only absolute numbers of anchor items or focusing on a single IRT model or equating method. Furthermore, previous research has primarily evaluated the absolute deviation of estimated parameters from true parameters. However, in diagnostic contexts, the correlation between these values is often more relevant for ensuring validity and test fairness. Therefore, the aim of this simulation study was to examine the impact of a broad range of key factors on test equating.
Methods:
We evaluated correlations and recovery indices between predefined true values and values estimated through test equating for three IRT parameters (discrimination, difficulty, and ability). To this end, we varied the equating method (MS, MM, MGM, IRF, TRF), the IRT model (2PL vs. 3PL), guessing probability (0.000-0.250), anchor item proportion (5-25%), test set size (20-80 items), and the discrimination parameters of the anchor items. In addition, we used samples of 25-100 individuals to assess equating quality under challenging conditions as well as samples of 500 and 1,000 individuals to reflect adequate modeling conditions.
Results:
Low guessing probabilities and high anchor item discrimination parameters strongly improved test equating quality for all three IRT parameters. Recovery of discrimination and ability parameters increased logarithmically with larger test set sizes and higher anchor item proportions, with each of these two factors partially compensating for reductions in the other. While sample sizes below 100 individuals produced inadequate parameter recovery, samples of 100 or 500 individuals were justifiable under certain conditions. However, samples of only 100 individuals carried a slight risk of non-convergence. The choice of the equating method had rather minor effects and the impact of the IRT model was ambivalent.
Discussion:
These findings highlight the importance of using distractor-free response formats without any guessing probability, anchor items with high discrimination parameters, and large samples to ensure valid test equating. For individual research and test application purposes, we provide a comprehensive data set covering multiple factor levels and a step-by-step simulation guide.
Related Concept Videos
The Nernst Equation
The interconnection between standard cell potentials and various thermodynamic parameters such as the standard free energy change ΔG° and equilibrium constant K has been previously explored. For example, a redox reaction involving zinc(II) and tin(II) ions at 1 M concentration with Eºcell = +0.291 V and ΔG° = −56.2 kJ is spontaneous.
Thermochemical Equations
Chemical Equations
Clausius-Clapeyron Equation
Henderson-Hasselbalch Equation
Nuclear Stability
To hold positively charged protons together...

