Related Experiment Video
Updated: Apr 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models in artificial intelligence to answer patient questions in spine surgery: an evaluation of
Janam Patel1, Zayaan Tirmizi1, Ayesha A Waheed1
1Department of Neurological Surgery, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Background:
Large language models (LLMs) are increasingly being explored in healthcare, particularly for enhancing patient education. In spine surgery, LLMs have the potential to enhance communication and support patients through perioperative care. However, concerns remain regarding the accuracy, readability, and overall reliability of these tools in delivering patient-facing information. This review aimed to understand the current use of LLMs in answering patient questions in spine surgery.
Methods:
A structured search of PubMed and Google Scholar was conducted using terms focused on LLMs and neurosurgery. Studies were only included if they tested LLMs' ability in answering patient questions related to spine surgery. Exclusion criteria included non-peer-reviewed articles, studies that did not evaluate chatbot performance, or those using LLMs for non-educational purposes.
Results:
LLMs were tested across a variety of spine-related topics, including scoliosis, lumbar and cervical fusion, endoscopic procedures, and spinal cord stimulation. Studies consistently reported moderate to high accuracy ratings. Readability scores remained a limitation, with most responses written at a college reading level. Empathy and clarity varied by model and condition, with some studies showing improved ratings when assessed by non-medical reviewers. Methodological variability across studies introduced inconsistencies and limited comparability.
Conclusions:
LLMs show promising utility for patient education in spine surgery for addressing frequently asked questions. However, challenges in readability, accuracy, and standardization limit their current clinical adoption. Moving forward, studies must incorporate standardized evaluation tools, address high rate of content hallucination, and focus on chatbot performance in personalized scenarios. Cross-disciplinary collaboration is essential to ensure safe, accessible integration into neurosurgical care pathways.

