Related Experiment Videos
Evaluating Generative Artificial Intelligence Chatbots in Answering Public Questions About Altered Consciousness: A
Introduction:
People increasingly use generative artificial intelligence chatbots to interpret health concerns before contacting clinicians. Although these systems are not formal autonomous emergency department triage tools, their responses may influence whether and when users seek professional care. Altered consciousness is especially high-risk because delayed or ambiguous escalation can postpone time-sensitive evaluation. This study evaluated the safety and quality of chatbot answers from an emergency nursing perspective.
Methods:
We conducted a cross-sectional comparative evaluation of ChatGPT, Gemini, Claude, DeepSeek, and Doubao. A 130-item candidate pool was screened to yield 66 public-facing English questions across 11 emergency consultation domains. Each question was submitted once to each chatbot in a fresh conversation between May 1 and May 30, 2026, producing 330 responses. Of note, 5 clinically experienced nurses independently rated the safety, accuracy, empathy, information quality, transparency, global quality, and readability of the responses.
Results:
Interrater agreement was good to excellent: Fleiss kappa for safety was 0.848, and intraclass correlation coefficients for nonbinary nurse-rated outcomes ranged from 0.836 to 0.878. Safe-response rates ranged from 89.4% for Doubao to 95.5% for ChatGPT, with no significant overall difference. ChatGPT had the highest mean accuracy and the highest scores on the DISCERN consumer-health-information instrument and the Journal of the American Medical Association benchmark criteria. Claude and Gemini were rated as the most empathic, and DeepSeek produced the most easily readable text across most readability formulas. Each chatbot generated at least 1 potentially unsafe answer.
Discussion:
Chatbots can support public education about altered consciousness, but omissions, unsafe sequencing, and urgency-softening make them unsuitable for autonomous prehospital or emergency department triage. Emergency nurses should help shape escalation-first, plain-language safety templates for high-risk chatbot advice.