Related Experiment Videos
Evaluating AI-Based Automated Essay Scoring Through Signal Detection Theory: Beyond Aggregate Agreement Metrics
1Methodology and Measurement, Australian Council for Educational Research, Melbourne, VIC, Australia.
Applied Psychological Measurement
|July 26, 2026
Summary
Large Language Models (LLMs) in educational assessment show lower accuracy than human raters. This study used Signal Detection Theory to find LLMs lack precision and systematically avoid high scores, impacting AI scoring system calibration.
Area of Science:
- Educational Measurement
- Artificial Intelligence
- Psychometrics
Background:
- Large Language Models (LLMs) are increasingly used in educational assessment, but current evaluation methods like Quadratic Weighted Kappa do not fully capture AI scoring nuances.
- Existing metrics obscure critical aspects such as discrimination accuracy and rater effects, limiting a deep understanding of AI performance in scoring.
Purpose of the Study:
- To apply Signal Detection Theory (SDT) for a detailed evaluation of state-of-the-art LLMs against expert human raters in essay scoring.
- To diagnose AI scoring behavior by decoupling evaluative discrimination from response criteria.
Main Methods:
- Utilized Signal Detection Theory (SDT) to analyze the scoring performance of eight advanced LLMs and expert human raters.
- Evaluated 1,726 essays, comparing AI and human discrimination accuracy and identifying systematic biases in AI scoring.
Main Results:
- Human raters demonstrated significantly higher evaluative precision, with discrimination estimates roughly double those of the evaluated LLMs.
- LLMs exhibited notable centrality effects and score compression, consistently failing to assign the highest scores according to rubric tiers.
- Discrepancies in human-machine agreement are attributed to both reduced discriminative accuracy in AI and shifts in response criteria.
Conclusions:
- The study provides a diagnostic framework for understanding and comparing AI scoring systems, moving beyond aggregate reliability metrics.
- Findings highlight the need for careful calibration and selection of AI scoring tools based on specific educational objectives and fairness considerations.