LLM-as-a-judge for infection prevention and control and antimicrobial resistance impact: comparing three main LLMs
Marcello Di Pumpo1,2, Leonardo Villani1,3, Maria Rosaria Gualano3
1Section of Hygiene, University Department of Life Science and Public Health, Università Cattolica del Sacro Cuore, Rome, Italy.
Background:
Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.
Methods:
We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence between automated and human assessments, adhering to CHART reporting guidelines.
Results:
Analysis of 404 evaluations revealed a systematic upward divergence: all LLMs consistently assigned higher scores than human experts. This optimism bias persisted after adjusting for domain-specific differences and clustering effects. The gap was particularly pronounced in domains of persuasiveness and AMR impact, while information quality showed more heterogeneous results. Intra-rater reliability assessments demonstrated that LLMs maintained stable scoring patterns under identical prompting conditions.
Conclusions:
LLMs exhibit a consistent leniency bias, systematically overestimating the quality of AMR-related health communication compared to human evaluators. These results do not support the use of LLMs for autonomous evaluation in high-stakes public health contexts. Rather, LLM-based judging is best suited as a scalable screening tool within supervised human-in-the-loop workflows, where expert oversight serves as a necessary safeguard for evidence-based accuracy.
Related Concept Videos
Healthcare Associated Infections II: Preventive Measures
The best practices for preventing healthcare-associated infections include hand hygiene, patient risk...
Clinical Significance of Antibiotic Resistance
Antimicrobial Effectiveness


