Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Response Surface Methodology01:16

Response Surface Methodology

288
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
288
Self-Report Tests of Personality01:22

Self-Report Tests of Personality

467
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
467
Multiple Regression01:25

Multiple Regression

3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Mechanistic Models: Compartment Models in Individual and Population Analysis01:23

Mechanistic Models: Compartment Models in Individual and Population Analysis

89
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
89
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

4.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.1K
Measures of Intelligence01:29

Measures of Intelligence

7.9K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.9K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The promise and challenges of computer mouse trajectories in DMHIs - A feasibility study on pre-treatment dropout predictions.

Internet interventions·2025
Same author

Changes in the Speed-Ability Relation Through Different Treatments of Rapid Guessing.

Educational and psychological measurement·2023
Same author

Disengaged response behavior when the response button is blocked: Evaluation of a micro-intervention.

Frontiers in psychology·2022
Same author

General cognitive ability assessment in the German National Cohort (NAKO) - The block-adaptive number series task.

The world journal of biological psychiatry : the official journal of the World Federation of Societies of Biological Psychiatry·2022
Same author

Erratum to: A Response-Time-Based Latent Response Mixture Model for Identifying and Modeling Careless and Insufficient Effort Responding in Survey Data.

Psychometrika·2022
Same author

A Response-Time-Based Latent Response Mixture Model for Identifying and Modeling Careless and Insufficient Effort Responding in Survey Data.

Psychometrika·2021

Related Experiment Video

Updated: Sep 19, 2025

Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

6.0K

Joint Item Response Models for Manual and Automatic Scores on Open-Ended Test Items.

Daniel Bengs1,2, Ulf Brefeld2, Ulf Kroehne1

  • 1Leibniz Institute for Research and Information in Education, Frankfurt, Germany.

Psychometrika
|June 16, 2025
PubMed
Summary

Automatic scoring of open-ended test items improves efficiency but introduces errors. New joint models accurately estimate student abilities by accounting for these automatic scoring errors, enhancing educational measurement.

Keywords:
automatic scoringitem response modelinglarge-scale assessment

More Related Videos

Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
07:43

Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios

Published on: August 4, 2023

2.2K
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.3K

Related Experiment Videos

Last Updated: Sep 19, 2025

Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

6.0K
Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
07:43

Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios

Published on: August 4, 2023

2.2K
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.3K

Area of Science:

  • Educational Measurement
  • Psychometrics
  • Artificial Intelligence in Education

Background:

  • Open-ended test items enhance construct validity but require costly manual scoring.
  • Manual scoring limits the use of item data in adaptive testing.
  • Automatic scoring using machine learning offers efficiency but introduces classification errors.

Purpose of the Study:

  • To develop statistical models that integrate both manual and automatic scores for open-ended items.
  • To account for classification errors inherent in automatic scoring within measurement models.
  • To improve the accuracy of ability estimation in educational assessments.

Main Methods:

  • Proposed two joint Item Response Theory (IRT) models incorporating both manual and automatic scores.
  • Extended existing IRT models to include a component for automatic scoring errors.
  • Evaluated models using data from the Programme for International Student Assessment (PISA) 2012 and simulated datasets.

Main Results:

  • The proposed joint models effectively mitigate the impact of classification errors on ability estimation.
  • Demonstrated improved accuracy compared to a baseline model that ignored automatic scoring errors.
  • Validated the models' performance on real-world (PISA) and simulated data.

Conclusions:

  • Joint modeling provides a robust approach to handling automatically scored open-ended items in educational testing.
  • Accounting for classification errors is crucial for accurate ability estimation when using automatic scoring.
  • These models enhance the validity and utility of open-ended items in adaptive and large-scale assessments.