Related Experiment Video
Updated: Jun 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of Large Language Models in Infectious Disease Decision-Making: From Examination to Clinical Practice
Dandan Wu1, Keying Chen1, Xiaocui Wu1
1Department of Infectious Diseases, The Second Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, Zhejiang, People's Republic of China.
Large language models (LLMs) show promise in infectious disease clinical reasoning, performing comparably to residents on exams. Human-AI collaboration enhances decision quality but clinician oversight remains crucial for safety.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Infectious Disease Management
Background:
- The integration of artificial intelligence (AI), specifically large language models (LLMs), into clinical practice is rapidly evolving.
- Evaluating the efficacy of LLMs in specialized medical domains like infectious diseases is critical for understanding their potential and limitations.
Purpose of the Study:
- To assess the performance of leading large language models (LLMs) in infectious disease clinical reasoning and decision-making.
- To compare LLM capabilities against infectious disease residents using a dual assessment framework.
Main Methods:
- A comprehensive evaluation of four LLMs (DeepSeek-R1, ChatGPT-5, Grok 3, Gemini 2.5 Flash) was performed.
- A dual assessment framework included examination-based questions and real-world clinical case scenarios.
- LLM performance was benchmarked against infectious disease residents, with outcomes measured by accuracy, score rates, and expert review using Likert scales.
Main Results:
- LLMs demonstrated comparable performance to infectious disease residents in examination-based assessments (p=0.54).
- LLMs excelled in low-order, knowledge-based questions, while residents showed an advantage in simple case-based questions requiring higher-order reasoning.
- LLM-assisted clinical decision-making showed trends toward improved accuracy and significantly enhanced completeness compared to independent decisions, though high-risk errors highlight limitations.
Conclusions:
- Large language models (LLMs) offer significant potential for infectious disease education and clinical decision support, particularly for knowledge recall.
- Clinician oversight is essential due to LLM limitations in complex clinical reasoning.
- A collaborative human-AI approach maximizes decision quality and safety, necessitating refined regulatory frameworks for responsible clinical deployment.
Related Concept Videos
Steps in Outbreak Investigation
Principles of Disease Surveillance
Introduction to Language of Pathophysiology ll
Investigation of Disease Outbreaks
Infectious Diseases and Their Occurrence
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
