Revisiting reliability with human and machine learning raters under scoring design and rater configuration in the
Xingyao Xiao1, Richard J Patz2, Mark R Wilson3
1Graduate School of Education, Stanford University, Stanford, California, USA.
The British Journal of Mathematical and Statistical Psychology
|January 31, 2026
Summary
Machine learning (ML) scoring offers a scalable alternative for constructed-response (CR) items, but human and ML rater bias significantly impacts ability estimation reliability. Careful scoring design and bias management are crucial for accurate results.
Area of Science:
- Educational measurement
- Psychometrics
- Artificial intelligence in education
Background:
- Constructed-response (CR) items assess higher-order skills but face challenges with human scoring variability and cost.
- Machine learning (ML) offers a scalable solution, but its psychometric implications in rater-mediated models require further investigation.
Purpose of the Study:
- To examine the impact of scoring design, rater bias, ML inconsistency, and model specification on ability estimation reliability in polytomous CR assessments.
- To compare the performance of different estimation models under various scoring conditions.
Main Methods:
- Monte Carlo simulation was used to manipulate human and ML rater bias, ML inconsistency, and scoring density (complete, overlapping, isolated).
- Five estimation models were compared, including the Partial Credit Model (PCM) and the Many-Facet Partial Credit Model (MFPCM).
Main Results:
- Systematic bias, rather than random inconsistency, was identified as the primary source of error in ability estimation.
- Hybrid human-ML scoring improved estimation with unbiased or opposing biases but exacerbated errors when biases aligned.
- The Partial Credit Model (PCM) with fixed thresholds consistently outperformed more complex models, and anchoring CR items stabilized MFPCM estimation.
Conclusions:
- Scoring design and bias structure, not model complexity, determine the effectiveness of hybrid scoring.
- Anchoring constructed-response items to selected-response metrics is a practical strategy for enhancing estimation stability in ML-assisted scoring.
Related Concept Videos
Reliability and Validity
13.9K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.9K
Controller Configurations
378
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
378
Electron Configurations
26.2K
Electron configurations and orbital diagrams can be determined by applying the Aufbau principle (each added electron occupies the subshell of lowest energy available), Pauli exclusion principle (no two electrons can have the same set of four quantum numbers), and Hund’s rule of maximum multiplicity (whenever possible, electrons retain unpaired spins in degenerate orbitals).
The relative energies of the subshells determine the order in which atomic orbitals are filled (1s, 2s, 2p, 3s, 3p,...
The relative energies of the subshells determine the order in which atomic orbitals are filled (1s, 2s, 2p, 3s, 3p,...
26.2K
Simplified Synchronous Machine Model
778
The Synchronous Machine Model is a fundamental tool in analyzing and ensuring the transient stability of power systems. This model simplifies the representation of a synchronous machine under balanced three-phase positive-sequence conditions, assuming constant excitation and ignoring losses and saturation. The model is pivotal for understanding the behavior of synchronous generators connected to a power grid, particularly during transient events.
In this model, each generator is connected to a...
In this model, each generator is connected to a...
778
Wind Turbine Machine Models
603
In the growing field of wind energy, incorporating wind turbine models into transient stability analysis is essential. Induction and synchronous machines are the primary models used, with induction machines being prevalent due to their simplicity and reliability.
Induction machines interact through the rotating magnetic field generated by the stator and the rotor. The key parameter is slip, which is the difference between synchronous speed and rotor speed relative to synchronous speed. Slip is...
Induction machines interact through the rotating magnetic field generated by the stator and the rotor. The key parameter is slip, which is the difference between synchronous speed and rotor speed relative to synchronous speed. Slip is...
603
Electron Configuration of Multielectron Atoms
64.9K
The alkali metal sodium (atomic number 11) has one more electron than the neon atom. This electron must go into the lowest-energy subshell available, the 3s orbital, giving a 1s22s22p63s1 configuration. The electrons occupying the outermost shell orbital(s) (highest value of n) are called valence electrons, and those occupying the inner shell orbitals are called core electrons. Since the core electron shells correspond to noble gas electron configurations, we can abbreviate electron...
64.9K


