Related Experiment Video
Updated: Jul 10, 2026

07:13
A Two-interval Forced-choice Task for Multisensory Comparisons
Published on: November 9, 2018
Comparing concurrent versus fixed parameter equating with common items: using the dichotomous and partial credit
Husein M Taherbhai1, Daer Yong Seo
1Harcourt Assessment, USA. husein_taherbhai@harcourt.com
Summary
Researchers found no significant benefit in choosing between fixed calibration and concurrent calibration methods for equating tests. Accuracy depends on factors like sample size and common items, not the calibration method itself.
Area of Science:
- Psychometrics
- Educational Measurement
- Statistical Modeling
Background:
- Equating is crucial for large-scale assessments to ensure score comparability across different test forms.
- Calibration methods, including anchoring, are debated for their impact on equating accuracy.
- Limited research exists comparing fixed calibration with concurrent calibration, particularly in psychometric practice.
Purpose of the Study:
- To compare the accuracy of fixed calibration versus concurrent calibration in test equating.
- To investigate the influence of sample size and common items on calibration method performance.
- To evaluate these methods using dichotomous Rasch and Rasch partial credit models.
Main Methods:
- Simulation study comparing fixed calibration and concurrent calibration.
- Utilized dichotomous Rasch and Rasch partial credit models.
- Employed the WINSTEPS computer program for calibration and equating.
Main Results:
- Contrary to expectations, concurrent calibration did not consistently yield greater accuracy in item parameter estimation.
- The accuracy of calibration methods was confounded by sample size and the number of common items.
- No definitive benefit was found for using one method over the other for parallel test forms.
Conclusions:
- The choice between fixed and concurrent calibration methods for equating parallel test forms offers no inherent advantage.
- Equating accuracy is influenced by design factors (sample size, common items) rather than the calibration process alone.
- Psychometricians should consider these confounding factors when selecting calibration strategies in large-scale assessments.
Related Concept Videos
Group Design
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between the two are due to...
Factorial Design
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Comparing Experimental Results: Student's t-Test
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
Theory of Attribution II: Kelley's Covariation Theory
Attribution theory plays a crucial role in social psychology, helping to explain how individuals interpret the causes of behavior. One prominent model within this field is Harold Kelley's covariation theory, which provides a systematic approach to determining whether internal traits or external circumstances drive a person's actions. The model posits that individuals rely on three key types of information—consensus, consistency, and distinctiveness—to make these judgments.Consensus: Comparing...
Comparison Tests
An infinite series composed of positive terms may either approach a finite value or increase without bound. Determining which outcome occurs is a central task in calculus, and comparison tests provide structured methods for making this determination. Rather than evaluating a series directly, these tests relate it to another series whose behavior is already known, allowing conclusions to be drawn through logical comparison.The direct comparison test applies to series with positive terms. If each...

