Related Experiment Video
Updated: Jun 26, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparison of artificial intelligence large language model chatbots in answering frequently asked questions in
Teresa P Nguyen1, Brendan Carvalho1, Hannah Sukhdeo1
1Department of Anesthesiology, Perioperative and Pain Medicine, Stanford School of Medicine, Stanford, CA, USA.
Background:
Patients are increasingly using artificial intelligence (AI) chatbots to seek answers to medical queries.
Methods:
Ten frequently asked questions in anaesthesia were posed to three AI chatbots: ChatGPT4 (OpenAI), Bard (Google), and Bing Chat (Microsoft). Each chatbot's answers were evaluated in a randomised, blinded order by five residency programme directors from 15 medical institutions in the USA. Three medical content quality categories (accuracy, comprehensiveness, safety) and three communication quality categories (understandability, empathy/respect, and ethics) were scored between 1 and 5 (1 representing worst, 5 representing best).
Results:
ChatGPT4 and Bard outperformed Bing Chat (median [inter-quartile range] scores: 4 [3-4], 4 [3-4], and 3 [2-4], respectively; P<0.001 with all metrics combined). All AI chatbots performed poorly in accuracy (score of ≥4 by 58%, 48%, and 36% of experts for ChatGPT4, Bard, and Bing Chat, respectively), comprehensiveness (score ≥4 by 42%, 30%, and 12% of experts for ChatGPT4, Bard, and Bing Chat, respectively), and safety (score ≥4 by 50%, 40%, and 28% of experts for ChatGPT4, Bard, and Bing Chat, respectively). Notably, answers from ChatGPT4, Bard, and Bing Chat differed statistically in comprehensiveness (ChatGPT4, 3 [2-4] vs Bing Chat, 2 [2-3], P<0.001; and Bard 3 [2-4] vs Bing Chat, 2 [2-3], P=0.002). All large language model chatbots performed well with no statistical difference for understandability (P=0.24), empathy (P=0.032), and ethics (P=0.465).
Conclusions:
In answering anaesthesia patient frequently asked questions, the chatbots perform well on communication metrics but are suboptimal for medical content metrics. Overall, ChatGPT4 and Bard were comparable to each other, both outperforming Bing Chat.
Related Concept Videos
Stages of General Anesthesia
General Anesthesia: Overview
General anesthesia induces unconsciousness in the whole body, while the others target specific areas or sensations. It is administered to minimize adverse effects, maintain...
Local Anesthetics: Common Agents and Their Applications
Cocaine is an ester of benzoic acid and methylecgogine. It is used to anesthetize and vasoconstrict locally. Currently, it is used primarily for topical applications. It is beneficial for surgeries on the upper respiratory tract, providing anesthesia and shrinking the mucosa. Cocaine in the form of cocaine hydrochloride is...
Local Anesthetics: Clinical Application as Spinal Anesthesia
Current Trends in Nursing II
Local Anesthetics: Clinical Application as Surface, Infiltration, and Conduction Block Anesthesia

