Related Experiment Video
Updated: Jun 12, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large-Scale Evaluation of Five Large Language Models in Anesthesia Decision-Making for Hip Fracture Surgery
Robert Chen1,2,3, Andrew Warburton1, Ron Do2
1From the Department of Anesthesiology, Perioperative and Pain Medicine.
Anesthesia and Analgesia
|June 10, 2026
Summary
Large language models (LLMs) offer potential for perioperative decision support but require careful evaluation. Their recommendations, while generally reasonable, may not align with current evidence, risking practice shifts without patient benefit.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Perioperative Medicine
Background:
- Large language models (LLMs) show promise for perioperative decision support.
- However, issues like hallucinations, miscalibration, and biases necessitate rigorous evaluation before clinical integration.
- Increasing LLM adoption in perioperative settings demands systematic assessment of their performance concerning patient and surgical factors.
Purpose of the Study:
- To systematically evaluate the performance of general-purpose and clinical large language models (LLMs) in perioperative decision support for hip fracture surgery.
- To determine how patient and surgical variables influence LLM recommendations.
- To assess the alignment of LLM-generated recommendations with contemporary clinical evidence.
Main Methods:
- Evaluated five general-purpose LLMs (DeepSeek 3.2, Gemini 2.5 Flash, GPT-5, GPT-5 mini, GPT-5 nano) using 216 hip fracture surgery vignettes.
- Generated 54,000 total responses across six surgery types, two sexes, and 18 patient variables.
- Used logistic regression to analyze effects on anesthesia type, peripheral nerve block placement, and arterial line placement; included sensitivity analysis with two clinical LLMs.
Main Results:
- All evaluated LLMs generally favored neuraxial over general anesthesia and appropriately recommended peripheral nerve blocks and arterial lines, with few sociodemographic biases.
- Free-text justifications often cited unsupported neuraxial benefits, and models issued strong recommendations lacking clear clinical justification.
- Clinical LLMs demonstrated similar recommendation patterns to general-purpose models.
Conclusions:
- Large language models (LLMs) provide generally sound perioperative recommendations but exhibit systematic preferences not fully supported by current evidence.
- Uncritical adoption of LLMs could lead to practice pattern shifts without guaranteed improvements in patient outcomes.
- Further research and critical evaluation are essential to ensure LLM integration enhances, rather than hinders, evidence-based perioperative care.
