Related Experiment Video
Updated: Sep 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models for tympanostomy patient education: readability and guideline adherence
Shreeya Bahethi1, Hetal Lad2, Shrey Shah2
1Department of Otolaryngology-Head and Neck Surgery, Hackensack University Medical Center, Hackensack, NJ, 07662, USA.
Purpose:
To evaluate the accuracy and readability of large language model (LLM)-generated patient education materials regarding tympanostomy tube placement.
Methods:
Over a two-month period, ChatGPT 4o, Gemini 2.5 Flash, and Google Search AI were prompted daily using long-form and layered prompt formats covering typical concerns regarding tympanostomy. Responses were scored on a 12-point rubric adapted from the AAO-HNS Clinical Practice Guidelines (CPG), assessing diagnostic accuracy, procedural clarity, and postoperative care. Readability was evaluated using Flesch Reading Ease and Flesch-Kincaid Grade Level. For each model and prompt type, average scores and variability were analyzed with 95% confidence intervals. Between-model differences were tested with Welch's t-tests and Cohen's d; temporal trends were analyzed via linear regression.
Results:
For long-form outputs, Google Search AI and Gemini 2.5 Flash demonstrated the highest mean CPG adherence (96.4% and 96.6%, respectively), both significantly exceeding ChatGPT 4o (84.3%; both P < .001; Cohen's d ≈ 2.05). For layered prompt sessions, Google Search AI again demonstrated the highest adherence (91.7%), followed by Gemini 2.5 Flash (88.1%) and ChatGPT 4o (76.2%). Guideline adherence was temporally stable across all models (P > .05 for most). All outputs exceeded the recommended sixth-grade reading threshold (mean FKGL, 9.4 for long-form; 10.6 for layered), with no statistically significant readability differences between prompt strategies.
Conclusions:
Google Search AI demonstrated the highest guideline concordance, though all models produced material too complex for typical patient comprehension. Structured prompting enhances clinical accuracy, but readability remains a key barrier to accessibility.