Related Experiment Video
Updated: Jun 7, 2026

Using Learning Outcome Measures to assess Doctoral Nursing Education
Published on: June 21, 2010
Evaluating the Quality of Artificial Intelligence-Generated Rubrics and Grading Reliability in Student-Written
Shannon Dicken1, Nicole DeLeon1, Eric J Gutierrez1
1University of Colorado Anschutz, Skaggs School of Pharmacy and Pharmaceutical Sciences, Aurora, CO, USA.
Generative artificial intelligence (genAI) can create grading rubrics and assess student reflections in pharmacy education. However, AI platform variability necessitates careful validation for reliable educational implementation.
Area of Science:
- Pharmacy education
- Artificial intelligence in education
Background:
- Pharmacy curricula increasingly utilize reflective writing assignments to assess student learning.
- Evaluating these assignments requires consistent and reliable grading rubrics.
Purpose of the Study:
- To evaluate the efficacy of generative artificial intelligence (genAI) in developing grading rubrics and grading student reflections in pharmacy education.
- To compare the performance of different genAI platforms (ChatGPT, Copilot, Gemini) in rubric generation and assignment grading.
Main Methods:
- Phase I: Three genAI platforms generated grading rubrics for student reflections. Faculty assessed rubric quality using interrater reliability.
- Phase II: Three genAI platforms and two human evaluators graded student reflections using a standardized rubric. Scoring consistency was analyzed using ANOVA and ICC.
Main Results:
- GenAI platforms produced rubrics of varying quality, with moderate interrater reliability among faculty.
- Significant differences were found in scores assigned by genAI platforms and human graders.
- Gemini showed the lowest agreement with other AI platforms and human graders.
Conclusions:
- GenAI tools can generate rubrics and grade assignments, showing potential for pharmacy education.
- Variability in genAI performance and inconsistent agreement with human graders highlight the need for rigorous validation prior to adoption.
- Further research is needed to optimize genAI use in educational assessment.
Related Concept Videos
Reliability and Validity
Self-Evaluation Maintenance Model
Social Foundations of Self III: Self-Evaluation
Self-Evaluation: Self-Enhancement and Self-Verification
Metacognition