Related Experiment Video
Updated: Aug 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Are large language models such as ChatGPT, capable of supporting patients and general practitioners after spine
Sebastian Wegmann1, Till Rosenkranz2, Philip Egenolf1
1Department of Orthopaedics, Trauma Surgery and Plastic and Aesthetic Surgery, University Hospital Cologne, Cologne, Germany.
Purpose:
To evaluate whether large language models (LLMs) can provide accurate, complete, and audience-adapted answers to common spine-surgery-related questions for patients and family practitioners.
Methods:
Ten frequently asked spine-surgery questions were collected at a level 1 trauma center and simplified linguistically. Five LLMs (ChatGPT, Claude 3.5 Sonnet, Gemini Advanced 1.5 Pro, Copilot Pro, and DeepSeek V3) were queried using zero-shot prompting with persona-specific instructions for family practitioners and middle-aged patients. Responses were assessed by spine surgeons and non-medical raters for correctness, completeness, adaptability, and empathy using five-point Likert scales. Readability was quantified using the Flesch Reading Ease Score (FRES).
Results:
All LLMs generated largely correct and usable responses. ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers. Gemini and Copilot achieved superior readability and empathy for patient-facing responses. DeepSeek demonstrated balanced performance across all domains. Readability differed substantially between practitioner- and patient-oriented outputs.
Conclusion:
LLMs can support communication and education following spine surgery when used with structured prompting. Clinical oversight remains essential to mitigate risks related to inaccuracies and hallucinations.
Level Of Evidence:
III.