Related Experiment Videos

Evaluating AI-Based Automated Essay Scoring Through Signal Detection Theory: Beyond Aggregate Agreement Metrics

Xiaoliang Zhou1

  • 1Methodology and Measurement, Australian Council for Educational Research, Melbourne, VIC, Australia.

Summary

Large Language Models (LLMs) in educational assessment show lower accuracy than human raters. This study used Signal Detection Theory to find LLMs lack precision and systematically avoid high scores, impacting AI scoring system calibration.

Related Concept Videos