Related Experiment Video
Updated: Oct 6, 2026

Oral Health Assessment by Lay Personnel for Older Adults
Published on: February 2, 2020
AI-Assisted Rubric-Based Essay Grading in Dental Hygiene Education: A retrospective observational study
Courtney D Routh1, Richard Halpin2, J Nathaniel Holland3
1Department of Periodontics and Dental Hygiene, UTHealth Houston, School of Dentistry, Houston, TX, USA Courtney.D.Routh@uth.tmc.edu.
Abstract:
Purpose Assessment of reflective and professionalism-focused writing is a critical component of health professions education; competencies such as ethical reasoning, communication, and self-awareness are foundational to professional identity formation and clinical readiness. The purpose of this study was to evaluate the consistency of a large language model (LLM) in rubric-based grading of reflective essays in dental hygiene education, compare agreement patterns and systematic differences between artificial intelligence (AI) and faculty grading, and examine its potential to support faculty-driven assessment processes.Methods Deidentified, archived reflective essays from first-year dental hygiene student cohorts (2022-2024; N = 90) were regraded using a secure, institutionally managed AI program (ChatGPT, Model 4o). Essays were originally assessed by faculty using a standardized rubric and re-evaluated by the LLM using identical criteria. Faculty and AI scores were compared overall and by domain using paired statistical tests. Grading time was recorded, and test-retest reliability was evaluated across three independent scoring runs.Results Domain-level analyses required complete rubric data across all five domains; because of incomplete rubric data, 77 submissions (N = 77) were included in domain-level analyses. Faculty assigned significantly higher overall scores than the AI system (p < 0.001), with an average difference of approximately 1.9 points on a 10-point scale. Significant differences were found in four of the five rubric domains. Artificial intelligence grading was rapid, averaging approximately 39 seconds per essay, and demonstrated strong internal consistency across repeated evaluations (ICC = 0.87). The AI scores showed greater dispersion and more conservative scoring patterns, whereas faculty scores demonstrated ceiling effects.Conclusion Artificial intelligence-based grading demonstrated strong reproducibility and rapid evaluation, with systematic differences in score calibration compared with faculty assessment. The use of AI may complement faculty-driven assessment by supporting consistent, scalable grading practices in dental hygiene education, potentially enhancing grading reliability while reducing faculty workload.