Related Experiment Video
Updated: May 11, 2026

Pedicle Screw Placement Using an Augmented Reality Head-Mounted Display in a Porcine Model
Published on: May 24, 2024
Prompt Engineering and Follow-Up Questioning Improves the Readability of Spine Surgery Questions in Large Language
Sohail Daulat1, Nikhil Dholaria2, Gregory Burnet1
1Department of Neurological Surgery, University of Pittsburgh School of Medicine, Pittsburgh, Pennsylvania, USA.
Background:
The field of spine surgery is complex, and patient education material within the field is often written at a reading level that exceeds the recommended standard. Large language models (LLMs), such as ChatGPT, have shown potential for generating educational content, but require further investigation to determine whether prompt-engineering or asking follow-up questions can improve readability. The purpose of this study is to evaluate which models, including newer versions, and prompting strategies have the greatest improvement in the readability.
Methods:
ChatGPT 4o and 5 were prompted with 45 standardized spine surgery questions across 5 common procedures. Each question underwent 5 prompting phases: baseline (phase 1), follow-up clarification (phase 1.5), sixth-grade level request (phase 2), rule-based prompting (phase 3), and direct readability targeting (phase 4). Readability was measured using Simple Measure of Gobbledygook, Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index, and Coleman-Liau Index scoring. Results were standardized and analyzed using paired t-test and Wilcoxon signed-rank test, post-hoc analysis, of variance, and interaction models. Furthermore, a resource for spine surgeons with the most readable answers was created.
Results:
While ChatGPT 4o generated significantly more readable responses than ChatGPT 5 across all phases except phase 1 (P < 0.001), ChatGPT 5 produced significantly more reliable citations (P < 0.05). Phase 2 yielded the most readable responses, with 51.11% meeting the sixth-grade level. Follow-up clarification questions and simplified prompt engineering were more effective than complex rule-based prompts.
Conclusions:
Prompt engineering and follow-up questioning significantly enhanced the readability of LLM-generated responses to spine surgery questions at a level appropriate for patients. Further investigation is needed to assess whether these responses are more readable in real-world clinical settings instead of relying on objective readability scoring methods.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy

