Related Experiment Video
Updated: Sep 4, 2026

Learning Modern Laryngeal Surgery in a Dissection Laboratory
Published on: March 18, 2020
Large Language Models as Examinees and Graders in Simulated General Surgery Oral Board-Style Cases: A Psychometric
Kian A Huang1, Haris K Choudhary1, Allan L Xu1
1General Surgery, University of South Florida Morsani College of Medicine, Tampa, USA.
Introduction:
Surgical oral board examinations are vulnerable to examiner variability, potentially compromising scoring reliability. This study evaluated six public frontier large language models (LLMs) as both examinees and graders on general surgery oral board-style cases.
Methods:
Six commercially available LLMs were assessed across three standardized cases using a fully crossed psychometric design. Responses were graded on a 0-3 ordinal scale by six LLMs and three blinded senior surgeons using a structured rubric with anchored descriptors for each score level. Analyses included performance comparisons, reliability testing, and variance decomposition.
Results:
All LLMs achieved passing or near-passing performance. Significant differences existed between models (Friedman χ² = 12.51, p = 0.028). AI graders demonstrated greater internal consistency than the three-surgeon human panel in this pilot study (Cronbach's α = 0.697 vs. -0.923), though this comparison is limited by the small human panel and the absence of rater calibration for human graders. The dominant source of score variance was the Examinee × Rater interaction, suggesting that grader-driven disagreement accounted for more score variability than actual examinee performance differences in this dataset.
Conclusions:
LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples. Future work should evaluate whether these findings generalize to live examination settings, compare against calibrated human rater panels, and explore AI-augmented panel designs.
