Related Experiment Videos
Readability of AI-Generated Patient Visit Summaries in Orthopedic Surgery: Retrospective Analysis
Edward Lee Major Iii1, Vivek P Shah2, Amber N Carroll3
1Department of Orthopaedic Surgery and Sports Medicine, College of Medicine, University of Kentucky, 740 S Limestone St, Lexington, KY, United States, 1 323 873 4940.
Background:
Patient visit summaries (PVS) are patient-facing documents intended to reinforce communication and promote patient education after clinical encounters. Despite national recommendations that patient education materials be written at or below a sixth-grade reading level, most orthopedic materials substantially exceed this threshold. Current visit summaries are also time-consuming to generate, lack personalization, and often fail to meet patient literacy needs. AI-based scribes can generate personalized PVS in real time directly from patient-provider conversations, offering a potential solution.
Objective:
This study aimed to evaluate the readability of AI-generated PVS produced by a commercial AI scribe platform in an orthopedic surgery setting and determine their alignment with established literacy standards for patient-facing materials.
Methods:
A total of 1007 consecutive AI-generated PVS from an academic orthopedic surgery outpatient clinic between December 2023 and May 2024 were reviewed. Following standardized preprocessing, including restoration of original section headings and removal of diagnosis label headers, summaries identified as incomplete were excluded (n=25), yielding a final study cohort of 982 summaries. Readability was assessed using 5 validated indices: the Flesch-Kincaid Grade Level (FKGL), Flesch Reading Ease Score (FRES), Gunning Fog Index (GFI), Coleman-Liau Index (CLI), and Simple Measure of Gobbledygook (SMOG) Index. The proportions of PVS meeting the sixth- and eighth-grade benchmarks were calculated. Spearman's rank correlation (ρ) assessed associations between word count and readability metrics. The Kendall Coefficient of Concordance (W) was used to evaluate agreement among indices after reverse-coding FRES for directional alignment. Statistical significance was set at P<.05.
Results:
Among 982 analyzed PVS, the mean FKGL was 9.3 (SD 1.2, 95% CI 9.2-9.4) and the mean FRES was 57.6 (SD 7.9, 95% CI 57.1-58.1), corresponding to "fairly difficult." Mean scores for secondary indices were 12.4 (SD 1.6) for GFI, 11.0 (SD 1.5) for CLI, and 12.4 (SD 1.2) for SMOG. Only 0.4% (4/982) of PVS met the sixth-grade benchmark and 14.2% (139/982) met the eighth-grade threshold. Word count was not significantly correlated with FKGL (ρ=0.012, P=.70), suggesting that sentence structure and vocabulary rather than document length drive reading complexity. Readability indices demonstrated strong agreement across all 5 metrics (W=0.882, P<.001).
Conclusions:
In this orthopedic outpatient setting, AI-generated PVS consistently exceeded recommended patient literacy thresholds, with a mean ninth-grade reading level and only 14.2% of summaries meeting the eighth-grade standard. In the absence of a concurrent comparison group, the present study was not designed to evaluate improvement over traditional methods. These findings establish a quantitative readability baseline for AI scribe output in orthopedic surgery and highlight the need for algorithmic refinement, plain language optimization, and prospective patient comprehension testing to improve health communication.