Related Experiment Video
Updated: Jun 11, 2026

Application of the Intelligent High-Throughput Antimicrobial Sensitivity Testing/Phage Screening System and Lar Index of Antimicrobial Resistance
Published on: July 21, 2023
An antibiotic chatbot: Evaluation of a retrieval-augmented generation approach for providing guideline-based
David W Eyre1, Ruth Corrigan2, Lauren Hookham2
1Big Data Institute, University of Oxford, Oxford, UK; National Institute for Health Research Health Protection Research Unit in Antimicrobial Resistance and Healthcare Associated Infection, University of Oxford, Oxford, UK; National Institute for Health Research Oxford Biomedical Research Centre, Oxford, UK; Oxford University Hospitals, Oxford, UK; Nuffield Department of Medicine, University of Oxford, Oxford, UK.
Background:
Large language models (LLMs) have potential to provide clinical infection advice, but variations in prevalent pathogens and antimicrobial resistance requires models to be adapted to local contexts. We evaluated a retrieval-augmented generation (RAG) approach to provide antibiotic and infection advice explicitly constrained to local guidelines.
Methods:
Relevant guideline sections from Oxford University Hospitals were identified combining keyword-matching and a medical embedding model. A locally-deployed LLM (gpt-oss-20b) generated answers using the retrieved context. Performance was assessed using 200 simulated questions with an LLM-as-judge, and 66 human-written questions reviewed by ≥2 infection specialists.
Results:
The model attempted to answer 186/200 (93%) simulated clinical advice queries, of which 162 (87%) responses were judged fully-correct, 14 (8%) partially-correct, and 10 (5%) incorrect. Performance was lower in complex scenarios, e.g., when renal impairment was present. For 57 human-written questions covered by guidelines, 46 (81%) single-stage responses were fully-correct and 10 (18%) partially-correct. Of 9 out-of-scope questions, 5 (56%) were correctly identified. A multi-stage pipeline modestly improved performance (84% fully-correct). Median answer generation time was 12 s (single-stage) and 15 s (multi-stage). LLMs without RAG-based local guideline context had lower performance: 21/186 (11%) answers to simulated questions fully correct with the same locally-deployed LLM and 92/200 (46%) with a current frontier model (gpt-5.4).
Conclusion:
An LLM grounded in local antimicrobial guidelines can deliver mostly accurate, concise infection advice but still generates occasional errors and does not always recognise out-of-scope queries. Further optimisation and safety mechanisms are required before routine clinical deployment.
More Related Videos
08:58Isolation and Identification of Waterborne Antibiotic-Resistant Bacteria and Molecular Characterization of their Antibiotic Resistance Genes
Published on: March 3, 2023
11:15Quadruple-Checkerboard: A Modification of the Three-Dimensional Checkerboard for Studying Drug Combinations
Published on: July 24, 2021
Related Concept Videos
Antibiotic Selection
Automated Microbial Diagnostics
Microbiota Modulation by Antibiotics
Antimicrobial Proteins
Interferons
Interferons (IFNs) are proteins produced by lymphocytes, macrophages, and fibroblasts infected with viruses. While IFNs cannot prevent viruses from entering and...
Microorganisms in Medicine and Therapeutics
Clinical Significance of Antibiotic Resistance