Related Experiment Video
Updated: Jun 4, 2026

13:12
Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
Large Language Model Hallucinations in Spine Surgery: A Comparative Analysis of Clinician vs Patient-Level Prompts
Nathan D McLaughlin1, Advika N Srinivas1, Zachary F Lowe1
1Division of Neurological Surgery, Department of Surgery, Saint Louis University School of Medicine, Saint Louis, Missouri, USA.
Neurosurgery Practice
|June 3, 2026
Summary
Large Language Models (LLMs) show significant citation errors in neurosurgery topics, especially for patient queries. Always verify LLM-generated medical information for patient safety.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Neurosurgery Research
Background:
- Large Language Models (LLMs) are increasingly used for medical information synthesis by professionals and patients.
- LLMs' tendency to generate fabricated information (hallucinations) poses a significant risk in healthcare.
- Citation accuracy of LLMs for neurosurgical spine topics requires quantification and comparison.
Purpose of the Study:
- To quantify and compare the citation accuracy of five prominent LLMs for common neurosurgical spine topics.
- To evaluate LLM performance on both clinician-level and patient-level queries.
- To identify differences in error profiles among various LLMs.
Main Methods:
- Five LLMs (ChatGPT, Gemini, Claude, Microsoft Copilot, OpenEvidence) were evaluated.
- Clinician-level and patient-level questions on spine surgery were posed to LLMs using specific persona prompts.
- Citations were manually verified and scored for accuracy (2=accurate, 1=misrepresented, 0=fabricated).
Main Results:
- OpenEvidence achieved perfect citation accuracy.
- Claude showed the highest accuracy among general LLMs (78.0%), while Gemini had the lowest (average score 0.97, 26.2% fabrications).
- General LLMs performed worse on patient queries, with increased fabrications and citation of non-peer-reviewed sources, especially for Copilot.
Conclusions:
- General-purpose LLMs exhibit substantial and variable citation errors for spine surgery topics, particularly for patient-facing prompts.
- Specialized platforms like OpenEvidence offer high accuracy, contrasting with widely accessible models.
- Careful verification of all LLM-generated medical information is crucial for patient safety.
