Related Experiment Video
Updated: Jun 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large-Scale Evaluation of Five Large Language Models in Anesthesia Decision-Making for Hip Fracture Surgery
Robert Chen1,2,3, Andrew Warburton1, Ron Do2
1From the Department of Anesthesiology, Perioperative and Pain Medicine.
Background:
Large language models (LLMs) show promise for perioperative decision support, but persistent issues, including hallucinations, miscalibration, and biases, indicate they require rigorous evaluation before clinical use. As LLM adoption increases in perioperative settings, systematic evaluation is needed to determine how patient and surgical factors affect performance.
Methods:
We evaluated five general-purpose LLMs (DeepSeek 3.2, Gemini 2.5 Flash, GPT-5, GPT-5 mini, GPT-5 nano) using 216 standardized hip fracture surgery vignettes crossing six surgery types, two sexes, and 18 patient variables. We generated 50 samples per combination for 54,000 total responses and collected both structured recommendations and free-text justifications. We used logistic regression to estimate effects on three primary outcomes: anesthesia type, peripheral nerve block placement, and arterial line placement. In a limited sensitivity analysis, we evaluated two clinical LLMs (OpenEvidence, Doximity GPT) with 36 responses each.
Results:
All models favored neuraxial over general anesthesia (76.1%-88.6% of responses), and all but DeepSeek 3.2 appropriately adjusted recommendations for relevant medical contraindications. All models except GPT-5 nano recommended preoperative peripheral nerve blocks (92.6%-99.3%) and were appropriately conservative regarding arterial line placement. However, free-text justifications frequently cited neuraxial benefits unsupported by recent randomized trials, and most models issued strong neuraxial recommendations despite a lack of clinical justification. We identified limited sociodemographic biases, with only one significant and clinically meaningful effect across 150 comparisons. Clinical LLMs provided similar recommendations to general-purpose models.
Conclusions:
While LLMs provided generally reasonable recommendations, systematic preferences diverging from contemporary evidence suggest uncritical use could shift practice patterns without improving patient outcomes.
