Related Experiment Video
Updated: May 5, 2026

The ITS2 Database
Published on: March 12, 2012
The Best of Two Worlds: IRT-Enhanced Automated Essay Interpretable Scoring
1Department of Educational Psychology, Faculty of Education, East China Normal University, Shanghai 200062, China.
Abstract:
The Automated Essay Scoring (AES) systems confront two fundamental challenges: opaque "black-box" decision-making that limits educator trust, and insufficient validation across linguistically diverse educational contexts. This study proposes IRT-AESF, an innovative framework that bridges educational measurement theory and artificial intelligence by integrating item response theory (IRT) with deep learning. The framework generates three theoretically grounded psychometric parameters: student ability, item difficulty, and item discrimination, which provide transparent and interpretable explanations for scoring decisions. We rigorously evaluated IRT-AESF through 5-fold cross-validation on three large-scale datasets comprising 41,328 authentic essays from English and Chinese educational settings, including both classroom assessments and high-stakes examinations. Results demonstrate statistically significant improvements over competitive baseline models, achieving an 8.4% relative increase in quadratic weighted kappa while maintaining robust cross-lingual performance. This research advances the development of transparent, trustworthy automated assessment systems that deliver not only scores but meaningful diagnostic insights for educational practice.
Related Concept Videos
Self-Evaluation: Self-Enhancement and Self-Verification
Automatic Processing and Automatic Social Behavior
Improving Translational Accuracy
Improving Translational Accuracy
Comparing Experimental Results: Student's t-Test