Related Experiment Video
Updated: Jan 6, 2026

Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
Performance of Large Language Models in Complex Anesthesia Decision-Making: A Comparative Study of Four LLMs in
Qian Ruan1, Jinghong Shi1, Yunke Dai1
1Department of Anesthesiology, Clinical Medical College, The First Affiliated Hospital of Chengdu Medical College, Chengdu, Sichuan, China.
Abstract:
To evaluate and compare the performance of four Large Language Models (LLMs) in anesthesia decision-making for critically ill obstetric and geriatric patients and analyze their decision reliability across different surgical specialties. Prospective comparative analysis using standardized case evaluations. Four LLMs (ChatGPT-4o, Claude 3.5 Sonnet, DeepSeek-R1, and Grok 3). Thirty complex surgical cases (10 obstetric, 20 geriatric; 8 specialties) were analyzed. A 12-dimensional framework tested the models using unified prompts and decision points. Five trained anesthesiologists independently evaluated the models across six dimensions (patient assessment, anesthesia plan, risk management, individualization, contingency planning, decision logic; 1-10 scale, total 6-60). Overall, DeepSeek performed best (51.43 ± 2.74 points), significantly outperforming other models (P < 0.001). For obstetric cases, the mean scores were: DeepSeek (52.00 ± 1.83), Grok (49.40 ± 3.06), ChatGPT (47.60 ± 2.88), and Claude (46.60 ± 2.17). For geriatric cases, scores were: DeepSeek (51.15 ± 3.10), Grok (48.60 ± 2.33), ChatGPT (47.35 ± 2.50), and Claude (45.75 ± 2.05). Across specialties, all models performed best in hepatobiliary surgery, burn surgery, and thoracic surgery. DeepSeek demonstrated consistent performance across all dimensions, with notable advantages in decision logic (8.80 ± 0.40) and contingency planning (8.27 ± 0.45). All LLMs demonstrated strong anesthesia decision-making capabilities, with DeepSeek showing the best overall performance. Exploratory analysis revealed performance variations across specialties, although small sample sizes preclude definitive conclusions. Clinical implementation should consider specialty-specific factors and decision process characteristics.
More Related Videos
Related Concept Videos
Stages of General Anesthesia
General Anesthesia: Overview
General anesthesia induces unconsciousness in the whole body, while the others target specific areas or sensations. It is administered to minimize adverse effects, maintain...
Local Anesthetics: Clinical Application as Spinal Anesthesia
Inhalational Anesthetics: Overview
Local Anesthetics: Clinical Application as Epidural Anesthesia
Since epidural anesthetics can be infused through an epidural catheter, all types of drugs, including short-acting ones, can be administered. Chloroprocaine and lidocaine are examples of short and long-duration anesthetics, respectively. Bupivacaine...
Parenteral Anesthetics: Overview

