Related Experiment Video
Updated: Oct 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Concordance of five large language models with infectious diseases physicians on clinical decision-making: a
Yakup Demir1, Hüsameddin Atay2, Eren Öztürk3
1Department of Infectious Diseases and Clinical Microbiology, Diyarbakır Memorial Hospital, Diyarbakır, Turkey. yakupdemir36@gmail.com.
Purpose:
To evaluate the accuracy of five large language models (LLMs) on a set of infectious diseases (ID) case-based multiple-choice questions and to compare their performance with that of practising ID physicians.
Methods:
A cross-sectional, web-based comparative survey was conducted among ID physicians and residents across Turkey. Ten original case scenarios (30 five-option multiple-choice questions; stratified by difficulty) were validated by a five-expert Delphi consensus panel (two rounds) using IDSA guidelines and EUCAST 2024 breakpoints as the reference standard. The same questions were put to five LLMs - Grok, ChatGPT-4o, Gemini 2.5, Claude Sonnet, and DeepSeek-V3 - via a standardised prompt in a single session. Primary outcomes were accuracy rate (%), Cohen's κ computed from observed response marginals, and an exact one-sided binomial test comparing each model's score against the observed physician-group accuracy, with Cohen's h as the effect-size measure for the difference between proportions.
Results:
Of 129 respondents, 112 were included; 17 were excluded for self-reported AI use during the case section (this exclusion criterion was empirically supported: excluded participants scored significantly higher than included participants, 68.8% vs. 61.6%, Mann-Whitney p = 0.0015). Physician accuracy was 61.6% (18.5/30; SD 8.2; κ = 0.41 relative to a chance rate of 20%, interpreted descriptively). Four of five LLMs showed statistically significant superiority over the physician group: Grok 90.0% (κ = 0.806; Cohen's h = 0.69; binomial p < 0.001), ChatGPT-4o 83.3% (κ = 0.682; h = 0.50; p = 0.009), Gemini 2.5 83.3% (κ = 0.700; h = 0.50; p = 0.009), Claude Sonnet 80.0% (κ = 0.646; h = 0.41; p = 0.026); DeepSeek-V3 (76.7%; κ = 0.573; h = 0.33) did not reach statistical significance (p = 0.062). In Turkey-specific endemic cases, all models scored 100% versus 73% for physicians. Two items showed low concordance with the guideline-derived reference in both physicians and all LLMs; post-hoc review indicated that one of these items (prosthetic valve endocarditis) reflected a timing-dependent guideline nuance not fully captured by the original vignette, rather than a clear guideline violation. No correlation was found between physician AI-trust scores and clinical accuracy (Spearman ρ=-0.008; p = 0.937).
Conclusion:
On this case-based multiple-choice assessment, most evaluated LLMs achieved higher guideline concordance than a physician comparison group, particularly for guideline-standard and regionally endemic scenarios. These findings should be interpreted as evidence of accuracy on a structured test format rather than of general clinical decision-making superiority, and do not establish that LLMs should be used unsupervised in real clinical practice, which typically involves collaborative, multidisciplinary input rather than unilateral decisions.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe and...
Investigation of Disease Outbreaks
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Factors Affecting Illness
For instance, risk factors are connected to illness, disability,...
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions