Related Experiment Video
Updated: May 6, 2026

Optimized Management of Endovascular Treatment for Acute Ischemic Stroke
Published on: January 18, 2018
Comparative performance of GPT-4 models and expert anaesthesiologists in peri-operative antithrombotic management: A
Antonio Pérez-Ferrer1, Elena Gredilla Díaz, Jesús de Vicente Sánchez
1From the Department of Anaesthesiology and Intensive Care, Infanta Sofía University Hospital (APF, RHG, RGG, CCO, MAC), Faculty of Medicine, Health and Sports, European University of Madrid (APF, RHG, CCO), Biomedical Research Foundation of Infanta Sofía and Henares University Hospitals (FIIB HUIS-HUHEN) (APF, RHG, RGG, CCO, MAC), Department of Anaesthesia and Intensive Care, La Paz University Hospital, Madrid, Spain (EGD, JDVS), Severo Ochoa University Hospital, Leganés (ARL), Ramón y Cajal University Hospital (MÁPR), Móstoles University Hospital, Madrid, Spain (YLB), San Carlos University Clinical Hospital, Madrid, Spain (SEDF).
Background And Objective:
Large language models based on transformer architecture are increasingly considered as clinical decision-support tools; however, evidence of their reliability compared with human experts in high-stakes peri-operative contexts remains limited. Peri-operative antithrombotic management requires precise, guideline-concordant decision-making. The objective of this study was to compare the performance of a general-purpose and a domain-specific, transformer-based language model with that of practising anaesthesiologists in peri-operative antithrombotic scenarios.
Methods:
A cross-sectional analytical study was conducted in a fully simulated, nonclinical environment across five university-affiliated anaesthesiology departments in Spain. Twenty-five hypothetical peri-operative scenarios involving anticoagulant or antiplatelet therapy were designed according to current evidence-based guidelines. Five anaesthesiologists generated responses to peer-created and self-authored scenarios (125 clinician responses). Two language models independently answered all scenarios. Completeness and accuracy were rated independently by three blinded experts.
Results:
The domain-specific model generated more complete responses (mean 4.52 ± 0.39) than anaesthesiologists (3.72 ± 0.43) and the general-purpose model (4.25 ± 0.48; P < 0.001). Accuracy did not differ significantly between groups ( P = 0.107), with the highest mean accuracy observed for the domain-specific model (4.31 ± 0.57). Inter-rater reliability ranged from 0.717 to 0.885. The domain-specific model produced no incomplete responses and 1.3% inaccurate answers, whereas clinicians used external resources in 77% of cases.
Conclusions:
In simulated peri-operative antithrombotic management scenarios, a domain-specific, transformer-based language model generated faster and more complete responses than clinician-generated answers, while accuracy was comparable across groups. These findings support further prospective evaluation of domain-adapted language models before clinical integration.
