Related Experiment Video
Updated: Sep 19, 2025

Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
Joint Item Response Models for Manual and Automatic Scores on Open-Ended Test Items
Daniel Bengs1,2, Ulf Brefeld2, Ulf Kroehne1
1Leibniz Institute for Research and Information in Education, Frankfurt, Germany.
Automatic scoring of open-ended test items improves efficiency but introduces errors. New joint models accurately estimate student abilities by accounting for these automatic scoring errors, enhancing educational measurement.
Area of Science:
- Educational Measurement
- Psychometrics
- Artificial Intelligence in Education
Background:
- Open-ended test items enhance construct validity but require costly manual scoring.
- Manual scoring limits the use of item data in adaptive testing.
- Automatic scoring using machine learning offers efficiency but introduces classification errors.
Purpose of the Study:
- To develop statistical models that integrate both manual and automatic scores for open-ended items.
- To account for classification errors inherent in automatic scoring within measurement models.
- To improve the accuracy of ability estimation in educational assessments.
Main Methods:
- Proposed two joint Item Response Theory (IRT) models incorporating both manual and automatic scores.
- Extended existing IRT models to include a component for automatic scoring errors.
- Evaluated models using data from the Programme for International Student Assessment (PISA) 2012 and simulated datasets.
Main Results:
- The proposed joint models effectively mitigate the impact of classification errors on ability estimation.
- Demonstrated improved accuracy compared to a baseline model that ignored automatic scoring errors.
- Validated the models' performance on real-world (PISA) and simulated data.
Conclusions:
- Joint modeling provides a robust approach to handling automatically scored open-ended items in educational testing.
- Accounting for classification errors is crucial for accurate ability estimation when using automatic scoring.
- These models enhance the validity and utility of open-ended items in adaptive and large-scale assessments.
More Related Videos
07:43Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
Published on: August 4, 2023
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Response Surface Methodology
The process of RSM involves several key steps:
Self-Report Tests of Personality
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Goodness-of-Fit Test
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...