Related Experiment Videos
Large Language Models in Perioperative Antithrombotic Management: A Blinded Scenario-Based Expert Evaluation
İrem Durmuş1, Merve Bulun Yediyıldız1
1Department of Anesthesiology and Reanimation, Kartal Dr. Lütfi Kırdar City Hospital, University of Health Sciences Turkey, 34865 Istanbul, Turkey.
Abstract:
Background: Perioperative antithrombotic management is a complex and high-risk aspect of anesthesia practice, requiring a careful balance between bleeding and thromboembolic risks. Large language models (LLMs) are increasingly explored for clinical decision support, but their reliability in complex perioperative antithrombotic scenarios remains uncertain. Methods: In this blinded, scenario-based study, 100 elective surgical cases, including grey-zone scenarios, were developed to reflect clinically relevant perioperative antithrombotic decisions. The European Society of Anaesthesiology and Intensive Care/European Society of Regional Anesthesia and Pain Therapy (ESAIC/ESRA) 2022 guideline was incorporated into the standardized prompt as a common framework for neuraxial safety and relevant antithrombotic interruption intervals. Four LLMs (GPT-5.4 Thinking, Claude Opus 4.5, DeepSeek v3.2 Reasoning and Qwen3.5-Plus) were evaluated using a standardized prompt. Model responses were presented without model identity and independently assessed by two experienced anesthesiologists, neither of whom was an author of the present study, across five clinical domains on a 5-point Likert scale, with a separate clinical applicability assessment. Results: Using a uniform four-domain composite across all 100 scenarios, overall expert-rated performance differed across models (Friedman χ2 = 28.75, p < 0.001; Kendall's W = 0.096). Claude had the highest median primary composite score, although its difference from DeepSeek was not statistically significant. Leave-one-domain-out sensitivity analyses showed that between-model separation was particularly sensitive to the clinical-rationale domain; exclusion of this domain reduced Kendall's W to 0.030. A secondary five-domain analysis restricted to the 51 common-proceed scenarios, all of which were standard rather than grey-zone cases, also showed an overall between-model difference (Friedman χ2 = 55.13, p < 0.001; Kendall's W = 0.360). Because each scenario was queried only once, fine-grained differences in model ordering should be considered provisional. Although all models performed well in drug discontinuation decisions, greater variability was observed in discontinuation timing, anesthesia choice and clinical reasoning. Postponement behaviour varied substantially across models, with DeepSeek and Qwen recommending postponement more often than ChatGPT and Claude. Conclusions: LLMs show promise as supportive tools in perioperative decision-making, but their performance remains variable, especially in complex situations. At present, they should be used cautiously and always under expert supervision, rather than as independent decision-makers.