Related Experiment Video
Updated: Aug 5, 2026

Robot-Assisted Transcanal Endoscopic Ear Surgery for Congenital Cholesteatoma
Published on: December 15, 2023
Comprehensive Evaluation of AI Consent Forms in Otolaryngologic Surgery
Sholem Hack1, Rebecca Attal1, Armin Farzad2
1City St. Georges University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center Ramat Gan Israel.
Introduction:
Surgical consent documents are frequently written at reading levels exceeding average health literacy. Large language models (LLMs) may offer a scalable approach to generating clearer, procedure-specific consent forms. This study evaluated the clarity, clinical accuracy, and acceptability of consent forms generated by GPT-4 and Claude for common otolaryngologic procedures.
Methods:
Twenty AI-generated consent forms (10 GPT-4.0, 10 Claude-2.1) were produced using standardized prompts. In a survey-based, non-clinical setting, five board-certified otolaryngologists independently rated each form for medical accuracy, readability, comprehensibility, legal/ethical sufficiency, and usability using a 4-point scale. A cross-sectional cohort of 300 English-speaking adults (15 raters per form) evaluated perceived clarity and signing comfort on 5-point Likert scales, and perceived trust using a binary (Yes/No) item, and completed eight binary quality assessments. A blinded subgroup (n = 10) compared AI-generated and official national health system templates across five Likert domains. Readability was assessed using Flesch-Kincaid Grade Level (FKGL).
Results:
Mean lay ratings for clarity across AI-generated forms were high overall. Claude demonstrated numerically higher scores than GPT-4 for clarity (4.72 vs. 4.68), perceived trust (reported as proportions), and signing comfort (4.40 vs. 4.27). However, when analyzed at the form level, differences between models were not statistically significant for clarity (mean difference 0.04; t(9) = -0.80; p = 0.44) or signing comfort (mean difference 0.13, t(9) = -1.68, p = 0.13). Across binary domains, ≥ 95% of participants affirmed adequate explanation of risks, benefits, and alternatives. Experts rated GPT-4 more accurate than Claude (2.2 vs. 1.5, p = 0.034). Mean Flesch-Kincaid Grade Level was lower for AI-generated forms compared to official templates. Although prompts targeted a 6th-8th grade reading level, achieved readability scores were slightly higher (8.8-9.4).
Conclusions:
In a non-clinical evaluation, AI-generated consent forms were perceived as clear and clinically complete, with model-specific trade-offs between perceived clarity and clinical detail. These perception-based findings-reflecting participant ratings of clarity, perceived trust, and willingness to sign rather than objective comprehension-are hypothesis-generating, and prospective clinical and legal validation in more representative patient populations is required.