Related Experiment Video
Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Evaluation and Performance of Large Language Models in Clinical Infection Control Scenarios: A Benchmark
Shuk-Ching Wong1,2,3, Edwin Kwan-Yeung Chiu3,4, Kelvin Hei-Yeung Chiu3,4
1Infection Control Team, Queen Mary Hospital, Hong Kong West Cluster, Hong Kong SAR, China.
Abstract:
Background: Infection prevention and control (IPC) in hospitals relies heavily on infection control nurses (ICNs) who manage complex consultations to prevent and control infections. This study evaluated large language models (LLMs) as artificial intelligence (AI) tools to support ICNs in IPC decision-making processes. Our goal is to enhance the efficiency of IPC practices while maintaining the highest standards of safety and accuracy. Methods: A cross-sectional benchmarking study at Queen Mary Hospital, Hong Kong assessed three LLMs-GPT-4.1, DeepSeek V3, and Gemini 2.5 Pro Exp-using 30 clinical infection control scenarios. Each model generated clarifying questions to understand the scenarios before providing IPC recommendations through two prompting methods: an open-ended inquiry and a structured template. Sixteen experts, including senior and junior ICNs and physicians, rated these responses on coherence, conciseness, usefulness and relevance, evidence quality, and actionability (1-10 scale). Quantitative and qualitative analyses assessed AI performance, reliability, and clinical applicability. Results: GPT-4.1 and DeepSeek V3 scored significantly higher on the composite quality scale, with adjusted means (95% CI) of 36.77 (33.98-39.57) and 36.25 (33.45-39.04), respectively, compared with Gemini 2.5 Pro Exp at 33.19 (30.39-35.99) (p < 0.001). GPT-4.1 led in evidence quality, usefulness, and relevance. Gemini 2.5 Pro Exp failed to generate responses in 50% of scenarios under structured prompt conditions. Structured prompting yielded significant improvements, primarily by enhancing evidence quality (p < 0.001). Evaluator background influenced scoring, with doctors rating outputs higher than nurses (38.83 vs. 32.06, p < 0.001). However, a qualitative review revealed critical deficiencies across all models, for example, tuberculosis treatment solely based on a positive acid-fast bacilli (AFB) smear without considering nontuberculous mycobacteria in DeepSeek V3 and providing an impractical and noncommittal response regarding the de-escalation of precautions for Candida auris in Gemini 2.5 Pro Exp. These errors highlight potential safety risks and limited real-world applicability, despite generally positive scores. Conclusions: While GPT-4.1 and DeepSeek V3 deliver useful IPC advice, they are not yet reliable for autonomous use. Critical errors in clinical judgment and practical applicability highlight that LLMs cannot replace the expertise of ICNs. These technologies should serve as adjunct tools to support, rather than automate, clinical decision-making.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Sputum Studies II: Culture and Sensitivity
Sputum culture and sensitivity is a medical procedure used to diagnose bacterial infections in the respiratory tract and select the most appropriate antibiotics for treatment. This process involves analyzing sputum samples of thick and opaque secretions produced in the lungs and airways. These samples are collected from patients and then sent to the laboratory for analysis.
The test can identify various pathogens responsible for respiratory infections, including Streptococcus,...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...

