Related Experiment Videos
Testing standards for AI-based scores in automated essay scoring
Rudolf Debelak1,2, Matthias Ziegler3
1Institute of Education, University of Zurich, Zurich, Switzerland.
Plos One
|July 31, 2026
Summary
Artificial intelligence (AI) models can score essays, but evaluating their psychometric quality requires specific methods. This study offers a framework to assess AI-based assessment validity, fairness, and reliability, finding a DistilBERT model robust yet showing potential fairness issues.
Area of Science:
- Psychometrics
- Artificial Intelligence
- Educational Measurement
Background:
- Large language models (LLMs) in AI and machine learning enable text evaluation for psychological and educational assessments.
- Automated essay scoring (AES) presents unique psychometric challenges compared to classical tests.
- Evaluating AI-generated scores requires robust frameworks for validity, fairness, and reliability.
Purpose of the Study:
- To develop and apply a standardized toolkit for evaluating the psychometric quality of AI-based assessments.
- To address challenges in scoring essays using AI models.
- To review and propose methods for assessing validity, fairness, and reliability in automated essay scoring.
Main Methods:
- Review of existing psychometric evaluation methods for AI-generated scores.
- Proposal of new methods for assessing validity, fairness, and reliability.
- Empirical evaluation using a DistilBERT model on the Hewlett Foundation dataset for automated essay scoring.
Main Results:
- The DistilBERT model demonstrated robust internal consistency (Spearman-Brown coefficients .77–.92).
- Empirical evidence supported the validity of the AI evaluation model.
- Potential fairness violations were indicated when comparing human and AI scores across different topics.
Conclusions:
- A standardized, replicable toolkit for evaluating AI-based assessment psychometric quality is provided.
- AI models show promise in automated essay scoring but require careful fairness evaluation.
- Researchers and practitioners can use this framework to ensure the quality of AI-driven assessments.
Related Concept Videos
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Reliability and Validity
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.