Related Experiment Video
Updated: May 5, 2026

16:17
The ITS2 Database
Published on: March 12, 2012
32.9K
The Best of Two Worlds: IRT-Enhanced Automated Essay Interpretable Scoring
1Department of Educational Psychology, Faculty of Education, East China Normal University, Shanghai 200062, China.
Behavioral Sciences (Basel, Switzerland)
|May 4, 2026
Summary
This study introduces IRT-AESF, a new framework combining item response theory (IRT) with AI for transparent automated essay scoring (AES). It enhances trust and provides diagnostic insights across diverse educational settings.
Area of Science:
- Artificial Intelligence in Education
- Educational Measurement
- Natural Language Processing
Background:
- Automated Essay Scoring (AES) systems face challenges with opaque decision-making and limited validation in diverse linguistic contexts.
- Educator trust in AES is hindered by the lack of transparency in scoring algorithms.
- Existing AES models often lack robust validation across different educational settings and languages.
Purpose of the Study:
- To propose IRT-AESF, an innovative framework integrating item response theory (IRT) with deep learning for transparent and interpretable automated essay scoring.
- To generate theoretically grounded psychometric parameters (student ability, item difficulty, item discrimination) for enhanced understanding of scoring decisions.
- To address the limitations of current AES systems by improving transparency and validation across diverse educational contexts.
Main Methods:
- Developed the IRT-AESF framework by integrating item response theory (IRT) principles with deep learning models.
- Generated three key psychometric parameters: student ability, item difficulty, and item discrimination.
- Evaluated the framework using 5-fold cross-validation on three large-scale datasets (41,328 essays) from English and Chinese educational settings.
Main Results:
- IRT-AESF demonstrated statistically significant improvements over baseline models.
- Achieved an 8.4% relative increase in quadratic weighted kappa, indicating enhanced scoring accuracy and reliability.
- Maintained robust cross-lingual performance, validating its effectiveness in diverse linguistic environments.
Conclusions:
- The IRT-AESF framework offers a transparent and interpretable approach to automated essay scoring.
- This research advances the development of trustworthy AI-powered assessment systems.
- The framework provides meaningful diagnostic insights beyond simple scores, supporting educational practice.
Related Concept Videos
Self-Evaluation: Self-Enhancement and Self-Verification
4.7K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
4.7K
Automatic Processing and Automatic Social Behavior
376
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
376
Improving Translational Accuracy
2.6K
2.6K
Improving Translational Accuracy
11.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.6K
Comparing Experimental Results: Student's t-Test
5.8K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
5.8K