Related Experiment Video
Updated: Jul 14, 2026

Biomechanical Changes Related to Low Back Pain: An Innovative Tool for Movement Pattern Assessment and Treatment Evaluation in Rehabilitation
Published on: December 13, 2024
ChatGPT improves readability in validated spine patient-reported outcome measures
George Abdelmalek1, Siraj Shaikh1,2, Daniel Coban1
1Department of Orthopedic Surgery, St. Joseph's University Medical Center, Paterson, NJ 07503, United States.
Background:
Spine patient-reported outcome measures (PROMs) frequently exceed recommended health literacy thresholds, limiting accessibility. Large language models (LLMs) such as ChatGPT can simplify medical text, but their effects on validated outcome instruments remain unclear.
Methods:
A cross-sectional analysis of validated spine PROMs was conducted. Seventy-seven PROMs identified in a prior readability analysis were revised using ChatGPT 4.0 through a standardized prompt instructing simplification to a sixth-grade reading level. Pre and postrevision readability metrics were assessed using Readable.com across multiple grade-level and linguistic indices. Revised PROMs were additionally evaluated for content fidelity using a predefined taxonomy assessing alterations in response scales, recall timeframes, and item meaning. Differences were analyzed using the Exact Sign Test (α = 0.05).
Results:
Eighteen of nineteen linguistic parameters improved significantly following ChatGPT revision (p < .05). Word count decreased by 18%, sentence complexity declined, and all readability indices improved (p < .001). About 7 of 9 grade-level metrics achieved NIH/AMA sixth-grade readability compliance following revision. However, 59.7% of PROMs contained at least one content-related error. The most common errors included alteration of validated response scales (23%), omission or simplification of recall timeframes (18%), and consolidation of multiple items into single prompts (16%).
Conclusions:
ChatGPT 4.0 substantially improved the readability of validated spine PROMs but frequently introduced structural modifications affecting validated content. Although LLMs may enhance linguistic accessibility, unsupervised PROM revision risks compromising measurement integrity. Structured implementation strategies incorporating expert review and psychometric validation may be necessary before AI-modified PROMs can be integrated into spine outcomes research.
Related Concept Videos
Methods of Documentation III: PIE
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care evaluation by...
