Longitudinal stability of IRT and equivalent-groups linear and equipercentile equating
Xiuyuan Zhang1, Paul A McDermott2, John W Fantuzzo2
1University of Pennsylvania, USA. xzhang@collegeboard.org
Psychological Reports
|December 18, 2013
Summary
Item Response Theory (IRT) equating showed fewer discrepancies between test forms A and B than linear or equipercentile methods for Head Start children. Further research is needed to fully understand the benefits of each equating approach.
Area of Science:
- Educational Measurement
- Psychometrics
- Developmental Psychology
Background:
- Accurate measurement of learning requires equivalent test forms.
- Equating methods are used to ensure scores from different test forms are comparable.
- Head Start programs serve young children, necessitating reliable assessment tools.
Purpose of the Study:
- To compare the effectiveness of three equating methods: common-item IRT equating, linear transformation, and equipercentile transformation.
- To evaluate these methods over an academic year using data from 1,667 Head Start children.
- To identify which equating method minimizes discrepancies between two presumably equivalent test forms (A and B).
Main Methods:
- A multiscale criterion-referenced test with two forms (A and B) was administered to 1,667 children over four points in an academic year.
- A randomly equivalent groups design was employed.
- Common-item IRT equating (concurrent calibration), linear transformation, and equipercentile transformation were used for equating.
Main Results:
- Item Response Theory (IRT) equating demonstrated different discrepancy patterns compared to linear and equipercentile methods over time.
- IRT equating resulted in marginally smaller mean score differences between forms A and B.
- IRT equating generated slightly fewer distributional discrepancies between the two test forms than linear and equipercentile equating.
Conclusions:
- IRT equating appears to offer a slight advantage in minimizing form-to-form score discrepancies and distributional differences.
- The results suggest mixed findings, indicating that no single method is universally superior.
- Additional research is necessary to fully elucidate the relative strengths and weaknesses of IRT, linear, and equipercentile equating methods in practice.
Related Concept Videos
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Routh-Hurwitz Criterion II
1.3K
In the application of the Routh-Hurwitz criterion, two specific scenarios can arise that complicate stability analysis.
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
1.3K
Rigid Body Equilibrium Problems - II
6.4K
A rigid body is in static equilibrium when the net force and the net torque acting on the system are equal to zero.
Consider two children sitting on a seesaw, which has negligible mass. The first child has a mass (m1) of 26 kg and sits at point A, which is 1.6 meters (r1) from the pivot point B; the second child has a mass (m2) of 32 kg and sits at point C. How far from the pivot point B should the second child sit (r2) to balance the seesaw?
Consider two children sitting on a seesaw, which has negligible mass. The first child has a mass (m1) of 26 kg and sits at point A, which is 1.6 meters (r1) from the pivot point B; the second child has a mass (m2) of 32 kg and sits at point C. How far from the pivot point B should the second child sit (r2) to balance the seesaw?
6.4K
Longitudinal Studies
706
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
706
Linear time-invariant Systems
1.1K
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
1.1K
Measures of Intelligence
13.0K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
13.0K


