Related Experiment Video
Updated: Aug 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Vocal Cord Dysfunction: Evaluating the Utility of AI Large Language Models for Patient Education
Kyle Cook1, Phil Tseng1, David Ahmadian2
1University of Arizona College of Medicine - Tucson Tucson Arizona USA.
Objective:
To compare provider preferences for patient education materials generated by OpenEvidence, ChatGPT 5 Extended Thinking, and a laryngologist with fellowship training in response to common patient questions about vocal cord dysfunction (VCD).
Study Design:
Cross-sectional survey study.
Methods:
A laryngologist compiled four common patient questions about VCD and authored expert responses. Equivalent responses for patients were generated using OpenEvidence and ChatGPT 5 Extended Thinking with a standardized role prompt, each in a new chat session, and all outputs were limited to four sentences. Responses were de-identified and presented in a REDCap survey. Forty-five healthcare providers ranked responses based on perceived medical accuracy, clarity, and ability to address patient concerns. Preferences were analyzed using chi-square goodness-of-fit and Friedman testing with post hoc pairwise comparisons.
Results:
Among the 45 respondents, the most represented specialties were otolaryngology (n = 17) and allergy/immunology (n = 12), and most were attending physicians (n = 31). ChatGPT 5 Extended Thinking received the most first choice selections for three of the four questions, whereas OpenEvidence and ChatGPT performed similarly on the definition question. Analyses of full rankings confirmed significant differences among sources for all four questions (all Friedman p ≤ 0.001): responses generated by AI were preferred over the laryngologist response for the definition, diagnosis, and treatment process questions, while ChatGPT was preferred over both comparators for the etiology question. Otolaryngology respondents were more likely than non-otolaryngology respondents to rank the laryngologist response first for the definition question only.
Conclusion:
Responses generated by AI were frequently preferred over a single expert comparator for common VCD questions, with ChatGPT 5 Extended Thinking performing best overall and OpenEvidence showing similar performance on several items. These findings support further evaluation of LLMs as adjunctive tools in laryngology patient education, while highlighting the need for clinician oversight, assessment of individual model performance, and studies of patient outcomes.
Related Concept Videos
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids, corniculates, and...
Introduction to Language of Pathophysiology ll
Introduction to Language of Pathophysiology l
