Related Experiment Video
Updated: May 11, 2026

Early Weight-Bearing Rehabilitation Protocol After Anterior Cruciate Ligament Reconstruction
Published on: March 1, 2024
Plain-language prompting improves readability of ChatGPT-generated patient education for meniscal surgery without
Aritra Chakraborty1, Drew Mattin2, Mitchell J Christiansen3
1Department of Orthopaedic Surgery and Rehabilitation, Loyola University Chicago Stritch School of Medicine, Maywood, IL 60153, USA.
Introduction/Objectives:
Large language models (LLMs) such as ChatGPT are increasingly used to generate patient education materials; however, default ChatGPT responses often exceed recommended readability levels set by the American Medical Association (AMA) and National Institutes of Health (NIH) health-literacy recommendations. The purpose of this study was to evaluate the readability and educational quality of ChatGPT-generated patient education on meniscal surgery and to determine whether a standardized plain-language prompt could improve readability without compromising accuracy, relevance, or depth.
Methods:
Sixteen standardized patient-focused questions regarding diagnosis, management, and prevention of meniscus tears were submitted to ChatGPT-5 and ChatGPT-4o, with three replicates per question to ensure standardization. Responses were assessed for accuracy against the American Academy of Orthopaedic Surgeons (AAOS) OrthoInfo and scored for relevance and depth using 5-point Likert scales. Readability was assessed using Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES). All baseline responses were subsequently rewritten using a plain-language prompt targeting a sixth- to eighth-grade reading level. Pre- and post-prompt readability metrics were compared using paired t-tests. Inter-rater reliability was measured with Cohen's kappa.
Results:
Both models demonstrated 100% factual accuracy across all baseline responses compared with OrthoInfo. Mean relevance and depth scores were high for ChatGPT-5 (4.49 ± 0.22; 4.39 ± 0.26) and ChatGPT-4o (4.55 ± 0.15; 4.53 ± 0.08). Baseline readability exceeded recommendations (FKGL 11.4-12.1; FRES 37-40). The plain-language prompt significantly improved readability for both models, reducing FKGL by approximately 5 grade levels and increasing FRES by 30-38 points (P < 0.001), with no loss of accuracy, relevance, or depth.
Conclusion:
ChatGPT generates accurate and relevant patient-directed content on meniscal surgery; however, readability frequently exceeded established health literacy standards. A simple, reproducible plain-language prompt reliably reduced reading level into the target range, offering a practical strategy for sports medicine surgeons to enhance informed consent discussions and patient education materials.
Level Of Evidence:
IV.
Related Concept Videos
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Peripheral Artery Disease V: Postoperative Nursing Management
Guidelines and Strategies for Safe Computer Charting
Maintain Confidentiality and Security:
Pre-Procedural Guidelines for Assessing Blood Pressure
Patient-centered Care
