Related Experiment Video
Updated: Apr 24, 2026

Emergency Undocking in Robotic Surgery: A Simulation Curriculum
Published on: May 20, 2018
Comparative evaluation of artificial intelligence chatbots for real-time guidance during intraoperative anesthetic
Abdallah Ahmed Mezel Al-Azzam1, Zaid Alkhateeb1, Ahmed Shahin1
1Department of Anesthesiology and Intensive Care, University of Jordan, Amman, Jordan.
Background:
Artificial intelligence (AI) chatbots are increasingly used in healthcare, but their ability to interpret anesthetic monitoring data during intraoperative crises remains unclear.
Objective:
To evaluate AI chatbots' responses to simulated anesthetic emergencies, with a focus on visual monitor interpretation compared to contextual case information.
Methods:
This simulation-based study was conducted using a high-fidelity patient monitor. Twenty intraoperative emergencies were designed as static monitor images with minimal clinical context. Five AI platforms, ChatGPT, Claude, Gemini, Copilot, and DeepSeek, were tested. Each scenario was submitted in a separate, new conversation with no advanced reasoning tools enabled. Six blinded evaluators scored 100 responses using the validated CLEAR tool. The primary outcome was the mean total score per chatbot; secondary outcomes included domain-specific scores and differences between visual and contextual scenarios.
Results:
ChatGPT achieved the highest mean score (4.4 ± 0.6), outperforming other platforms across all CLEAR domains (P < 0.001). Its accuracy was consistent between contextual (4.6 ± 0.3) and visual (4.1 ± 0.9) scenarios. DeepSeek scored the lowest overall score (2.7 ± 1.1), and had a mean 2.3 ± 1.1 in "lack of false information", often due to misinterpreting monitor values, reducing its visual scenario performance. Overall, contextual scenarios scores were higher than visual ones across all platforms.
Conclusion:
Among AI chatbots, ChatGPT demonstrated the most consistent and guideline-concordant responses to simulated anesthetic crises as of March 2025. This study provides a benchmark for evaluating clinical AI performance and supports the selective integration of such tools in anesthesia decision-making workflows.
