Related Experiment Video
Updated: Aug 24, 2026

E-Patient Counseling Trial (E-PACO): Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Large language models versus expert clinicians in optimizing Chinese patient education materials for
Jing Han1, Yan-Li Shi2, Li Dai2
1Department of Stomatology, Jinan Maternal and Child Health Care Hospital, Jinan, China.
Objective:
To evaluate the performance of large language models (LLMs) and expert clinicians in optimizing Chinese patient education materials (PEMs) for temporomandibular disorders (TMD) across readability, accuracy, actionability, and cultural adaptability, and to determine whether a human-AI collaboration model can achieve an optimal balance among these dimensions.
Methods:
Fifteen TMD education topics were selected through a Delphi consensus process. Four groups of PEMs were generated: original texts (Group A), LLM-rewritten texts (Group B, Claude Opus 4.6), clinician-rewritten texts (Group C), and human-AI collaboration texts (Group D). Blinded assessments were conducted by a professional panel (n = 3) and a health literacy-stratified patient panel (n = 6). Readability was measured by sentence length and common character proportion, accuracy by 5-point expert ratings against standardized checklists, actionability by the PEMAT instrument, and cultural adaptability by qualitative thematic analysis. Prompt sensitivity was tested across three distinct styles.
Results:
LLM-rewritten texts demonstrated significantly shorter sentences (21.1 vs. 33.9 characters, p < 0.001) and substantially higher actionability (95.6% vs. 16.7%, p < 0.001) compared with original texts. Clinician-rewritten texts achieved the highest accuracy (4.48 vs. 3.23, p < 0.001) but low actionability (46.7%). The human-AI collaboration model matched clinician accuracy (4.49, p = 1.000) while preserving LLM actionability (95.6%), with 72% less clinician time. Sensitivity analysis confirmed actionability robustness under empathetic and authoritative prompts but attenuation under a concise prompt. Qualitative analysis identified four cultural adaptability themes, with LLMs excelling in terminology localization and behavioral structuring but showing limitations in TCM conceptual depth.
Conclusion:
The human-AI collaboration model represents a promising and balanced approach for producing Chinese TMD patient education materials, combining LLM strengths in structural optimization with expert clinician oversight for accuracy and cultural appropriateness.
