Related Experiment Video
Updated: Aug 21, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models in Spine Surgery: A Scoping Review of Clinical Efficacy, Technical Integration, and Ethical
Samer G Salman1,2, Rohan A Phadke1,2, Anne E Tatooles1
1School of Medicine, Baylor College of Medicine, Houston, TX, United States.
None:
Study DesignScoping review.ObjectivesTo map spine literature on large language models, characterize reported use cases, and identify evidence gaps limiting implementation.MethodsA scoping review was conducted according to Joanna Briggs Institute methodology and PRISMA-ScR guidance. PubMed, Embase, Scopus, Web of Science, and Cochrane were searched for English-language, peer-reviewed studies published from January 2023 through May 2026 that evaluated large language models in spinal disease, spine surgery, or spine-related care. Eligible studies were synthesized across clinical decision support, triage, patient communication, automation, surgical education, and implementation barriers.ResultsFifteen studies met inclusion criteria. Most evidence involved early evaluation of commercially available or general-purpose models rather than prospectively validated spine-specific systems. Reported applications included patient education, report simplification, coding support, emergency consultation simulation, spinal cord stimulation referral screening, conservative triage, and surgical education. Performance was strongest for structured text-based tasks, patient communication, documentation support, and simplified decision pathways. Performance was weaker for image interpretation, quantitative radiographic assessment, individualized operative planning, and granular procedure selection. Recurrent limitations included hallucinated or unsupported outputs, unreliable citation generation, limited multimodal capability, privacy and data-governance concerns, bias, unclear medicolegal accountability, and minimal validation.ConclusionsLarge language models are an adjunct in spine surgery, with the near-term role in clinician-supervised, text-centered workflows including patient communication, education, documentation, coding, guideline retrieval, and preliminary triage. Current evidence does not support autonomous diagnostic, radiographic, or operative decision-making. Future studies should prioritize spine-specific retrieval-augmented systems, validated multimodal workflows, privacy-preserving deployment, fairness assessment, and prospective evaluation using clinically meaningful outcomes.