Related Experiment Video
Updated: May 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Diagnostic accuracy and citation integrity of four large language models on otolaryngology vignettes
William Pennington-FitzGerald1, Akshay Warrier2, Sally Durant2
1Department of Otolaryngology-Head and Neck Surgery, Boston University Medical Center, One Boston Medical Center Pl, Boston, MA, 02118, USA. william.pennington-fitzgerald@bmc.org.
Objectives:
This study aimed to compare the diagnostic accuracy and citation integrity of four large language models (LLMs) including one general (ChatGPT-4) and three intended for clinical and research use (OpenEvidence, Perplexity, and Pathway), using standardized otolaryngology clinical vignettes.
Methods:
One hundred validated otolaryngology clinical vignettes were presented to each LLM with a prompt requesting both a diagnosis and supporting citations. Diagnostic accuracy was determined against reference answers, and errors were categorized as logical, informational, or explicit. Citation number, source type, hallucination rate, and journal CiteScores were also compared.
Results:
All models demonstrated high diagnostic accuracy (82.0%-91.0%), with ChatGPT-4 achieving the highest numerical accuracy (91.0%), though differences between models were not statistically significant (p = 0.057). Logical errors were most frequent across all models. OpenEvidence and Perplexity generated the most citations per response, while ChatGPT-4 produced the fewest and had the highest hallucination rate (23.0%). Source preferences varied, with OpenEvidence and Pathway favoring narrative reviews and Perplexity favoring government/public health websites. OpenEvidence had the highest mean CiteScore for journal citations.
Conclusion:
This is the first study to assess both diagnostic accuracy and citation integrity of LLMs in otolaryngology. While ChatGPT-4 was most accurate, it had the highest rate of citation hallucinations, suggesting a trade-off between accuracy and source reliability. OpenEvidence, though slightly less accurate, provided more consistent and verifiable references, demonstrating the ability to prioritize citation integrity alongside diagnostic performance for clinical integration.
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Improving Translational Accuracy
Improving Translational Accuracy
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Introduction to Language of Pathophysiology ll
Introduction to Language of Pathophysiology l
