Related Experiment Video
Updated: Jan 31, 2026

Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
A Comparison of Ten Large Language Models and a Conventional Search Engine for Clinical Decision Support in
Oliver Cafferty1,2, Sean D Jeffries1,2, Eric D Pelletier1,2
1From the Department of Surgical Interventional Sciences, McGill University, Montreal, Canada.
Background:
Advances in artificial intelligence (AI) have enabled large language models (LLMs) to generate complex and contextually relevant medical responses. However, their potential in clinical decision support for anesthesiology remains underexplored. This study evaluated the accuracy and clinical relevance of high-performing LLMs in response to anesthesia-related questions and compared their performance with traditional online search methods. Clinician perceptions of AI were also assessed. We hypothesized that top-performing large language models would outperform lower-tier models and traditional internet search tools by generating responses rated as more accurate, complete, and clinically relevant to anesthesiology-focused questions, as measured by higher mean evaluator scores on a 10-point Likert scale.
Methods:
Ten LLMs: GPT-4o, Claude-Sonnet 3.5, DeepSeek R1, Llama 3.1 Instruct 70B, Gemini 2.0, GPT o1-preview, GPT o1, GPT o3-mini, NOVA Pro, and Mistral. All models were tested using ten common general anesthesia questions developed by TMH and validated by 6 physicians. Two Google search conditions served as baselines: a default search conducted in a cleared browser (unpersonalized), and a personalized Google Snippet Search performed in a browser regularly used by a clinician. Four board-certified anesthesiologists independently rated each response on a 10-point Likert scale. An ad hoc Physician Perception Questionnaire captured clinicians' use of AI, trust in its output, and reliance on traditional information sources.
Results:
LLM performance varied significantly (F = 5.89, P < .0001). DeepSeek R1 achieved the highest overall score (7.7), whereas Gemini 2.0 Flash recorded the lowest among LLMs (5.2). The Google Snippet Search scored 5.3, the lowest overall. Pairwise Welch's t tests showed that DeepSeek R1 significantly outperformed Llama, o3-mini, and Mistral ( P < .001). Survey results indicated limited AI use in clinical practice; clinicians prioritized source credibility and continued to favor traditional resources.
Conclusions:
Although LLM-generated responses differed in quality, DeepSeek R1 and Claude-Sonnet 3.5 produced answers most consistent with expert clinical judgment. The poor performance of several models, coupled with clinician skepticism, underscores the need for further validation before integrating AI into routine anesthesiology decision support.
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
The Sense of Self: Reflected Self-Appraisal and Social Comparison
Subliminal Perception
Factors Affecting Perception
An illustrative example of a perceptual set is the scenario where an airline pilot told...
Perception
Bottom-up processing begins at the sensory level, where receptors detect external environmental stimuli. These could include the tactile sensation of...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.

