Related Experiment Video
Updated: Mar 21, 2026

E-Patient Counseling Trial E-PACO: Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Evaluating the Effectiveness of ChatGPT Versus Human Proctors in Grading Medical Students' Post-OSCE Notes
Kirstyn Thomas1, Laura Szalacha2, Karim Hanna3
1Family Medicine, University of South Florida/BayCare Health System, Tampa, FL, United States.
Background And Objectives:
Artificial intelligence (AI) tools have potential utility in multiple domains, including medical education. However, educators have yet to evaluate AI's assessment of medical students' clinical reasoning as evidenced in note-writing. This study compares ChatGPT with a human proctor's grading of medical students' notes.
Methods:
A total of 127 subjective, objective, assessment, and plan notes, derived from an objective structured clinical examination, were previously graded by a physician proctor across four categories: history, physical exam, differential diagnosis/thought process, and treatment plan. ChatGPT-4, using the same rubric, was tasked with evaluating these 127 notes. We compared AI-generated scores with proctors' scores using t tests and χ2 analysis.
Results:
The grades assigned by ChatGPT were significantly different than those assigned by proctors in history (P<.001), differential diagnosis/thought process (P<.001), and treatment plan (P<.001). Cohen's d was the largest for treatment plan at 1.25. The differences led to a significant difference in students' mean cumulative grade (proctor 23.13 [SD=2.84], ChatGPT 24.11 [SD 1.27], P<.001), affecting final grade distribution (P<.001). With proctor grading, 81 of the 127 (63.8%) notes were honors and 46 of the 127 (36.2%) were pass. ChatGPT gave significantly more honors (118/127 [92.9%]) than pass (9/127 [7.1%]).
Conclusions:
When compared to a human proctor, ChatGPT-4 assigned statistically different grades to students' SOAP notes, although the practical difference was small. The most substantial grading discrepancy occurred in the treatment plan. Despite the slight numerical difference, ChatGPT assigned significantly more honors grades. Medical educators should therefore investigate a large language model's performance characteristics in their local grading framework before using AI to augment grading of summative, written assessments.
More Related Videos
07:32Use of Galvanic Skin Responses, Salivary Biomarkers, and Self-reports to Assess Undergraduate Student Performance During a Laboratory Exam Activity
Published on: February 10, 2016
09:52Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide
Published on: January 15, 2017
Related Concept Videos
Methods of Documentation II: POMR
Blind Procedures