Related Experiment Video
Updated: Mar 21, 2026

E-Patient Counseling Trial E-PACO: Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Evaluating the Effectiveness of ChatGPT Versus Human Proctors in Grading Medical Students' Post-OSCE Notes
Kirstyn Thomas1, Laura Szalacha2, Karim Hanna3
1Family Medicine, University of South Florida/BayCare Health System, Tampa, FL, United States.
Artificial intelligence (AI) grading of medical student notes differs from human proctors, particularly in treatment plans. While statistically significant, the practical differences in AI assessment of clinical reasoning are small.
Area of Science:
- Medical Education Technology
- Artificial Intelligence in Healthcare
Background:
- Artificial intelligence (AI) tools show promise in various fields, including medical education.
- The utility of AI for assessing medical students' clinical reasoning in written notes remains underexplored.
Purpose of the Study:
- This study aimed to compare the grading of medical students' clinical notes by ChatGPT-4 against evaluations by a human proctor.
Main Methods:
- 127 subjective, objective, assessment, and plan (SOAP) notes from an objective structured clinical examination were analyzed.
- ChatGPT-4 and a human proctor used the same grading rubric to evaluate notes across history, physical exam, differential diagnosis, and treatment plan.
- Statistical analyses included t tests and chi-squared analysis to compare scores.
Main Results:
- ChatGPT-4 assigned significantly different grades compared to proctors in history, differential diagnosis, and treatment plan (P<.001).
- The largest discrepancy was observed in the treatment plan (Cohen's d=1.25).
- ChatGPT-4 resulted in a higher mean cumulative grade and a greater proportion of honors grades (92.9%) compared to proctor grading (63.8%).
Conclusions:
- ChatGPT-4's grading of student SOAP notes showed statistically significant differences from human proctor evaluations, with notable variations in treatment plan assessment.
- Despite numerical differences, the practical impact on overall grades was small, but AI assigned significantly more honors.
- Medical educators must carefully evaluate AI performance within their specific grading frameworks before implementing AI for summative assessments.
More Related Videos
07:32Use of Galvanic Skin Responses, Salivary Biomarkers, and Self-reports to Assess Undergraduate Student Performance During a Laboratory Exam Activity
Published on: February 10, 2016
09:52Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide
Published on: January 15, 2017
Related Concept Videos
Methods of Documentation II: POMR
Blind Procedures