Related Experiment Video
Updated: Apr 24, 2026

06:48
Emergency Undocking in Robotic Surgery: A Simulation Curriculum
Published on: May 20, 2018
9.6K
Comparative evaluation of artificial intelligence chatbots for real-time guidance during intraoperative anesthetic
Abdallah Ahmed Mezel Al-Azzam1, Zaid Alkhateeb1, Ahmed Shahin1
1Department of Anesthesiology and Intensive Care, University of Jordan, Amman, Jordan.
Saudi Journal of Anaesthesia
|April 23, 2026
Summary
ChatGPT demonstrated the most consistent responses in simulated anesthetic emergencies. This study benchmarks AI performance for safe integration into anesthesia decision-making.
Area of Science:
- Anesthesiology
- Artificial Intelligence
- Medical Simulation
Background:
- Artificial intelligence (AI) chatbots are increasingly utilized in healthcare settings.
- The efficacy of AI chatbots in interpreting anesthetic monitoring data during intraoperative crises is not well-established.
Purpose of the Study:
- To evaluate the performance of AI chatbots in responding to simulated anesthetic emergencies.
- To compare AI chatbot interpretation of visual anesthetic monitoring data versus contextual case information.
Main Methods:
- A simulation-based study using a high-fidelity patient monitor with 20 designed intraoperative emergencies presented as static images.
- Five AI platforms (ChatGPT, Claude, Gemini, Copilot, DeepSeek) were tested, with responses scored by six blinded evaluators using the CLEAR tool.
- Evaluations focused on accuracy, guideline concordance, and the absence of false information, comparing visual-only versus context-inclusive scenarios.
Main Results:
- ChatGPT achieved the highest mean score, outperforming all other platforms across all evaluation domains (P < 0.001).
- ChatGPT's performance was consistent across both contextual and visual-only scenarios.
- DeepSeek scored lowest overall, with particular difficulty in interpreting monitor values, impacting its visual scenario performance.
Conclusions:
- ChatGPT exhibited the most consistent and guideline-concordant responses to simulated anesthetic crises as of March 2025.
- This research establishes a benchmark for assessing clinical AI performance in anesthesiology.
- Findings support the careful integration of AI tools into anesthesia decision-making workflows.
