Related Experiment Video
Updated: Aug 24, 2026

Multiplex Therapeutic Drug Monitoring by Isotope-dilution HPLC-MS/MS of Antibiotics in Critical Illnesses
Published on: August 30, 2018
Testing Knowledge Boundaries: Adversarial Evaluation of LLMs for Antimicrobial Stewardship
Ángela Abejez-Arrizabalaga1, Galadriel Pellejero-Sagastizabal2, Rocío Aznar-Gimeno3
1Division of Infectious Diseases, Hospital Clínico Universitario Lozano Blesa, Zaragoza, Spain; Instituto de Investigación Sanitaria Aragón, Zaragoza, Spain; Centro de Investigación Biomédica en Red, CIBERINFEC, Madrid, Spain.
Objectives:
Evaluate whether general-purpose large language models (LLMs) demonstrate competencies suitable for antimicrobial stewardship (AMS) support and characterize their failure modes.
Methods:
Cross-sectional evaluation of seven LLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, Llama-3.3-70b-instruct, Qwen 2.5-72b-instruct, DeepSeek-chat-v3.1) using 30 clinical scenarios mapped to ESCMID AMS competency frameworks. Scenarios included deliberate traps for fabrication and dangerous recommendations. Six AMS experts from the Netherlands and Spain performed blinded dual evaluation using content scores (0-5 scale) and binary safety flags for fabrication and danger. Standard and incentivizing prompt framings were compared.
Results:
Four commercial models achieved mean content scores above 3.9/5.0: Claude Sonnet 4.5 (4.06), Gemini 2.5 Pro (3.96), Grok 4 (3.96), and GPT-5 (3.94). Open-weight models scored significantly lower (2.94-3.57). No model achieved more than 63% responses free of fabrication or danger flags. However, fabrication did not impair clinical utility in non-trap scenarios (all within-category comparisons p>0.20). Danger flags ranged from 6.7% to 16.7% across models, with no significant difference between commercial and open-weight models. Incentivizing prompts were associated with a consistent 0.48-point-content score improvement (p=0.006), though significance attenuated after accounting for scenario-level clustering. Evaluators endorsed LLMs as useful AMS support tools with moderate supervision (5/6), identifying documentation preparation and trainee education as promising applications.
Conclusions:
Medically untrained LLMs demonstrate competencies suitable for supervised AMS support. Fabrication remains the central safety challenge and requires verification workflows; danger, though less frequent (6.7-16.7%), concentrated in identifiable and therefore mitigable failure modes. Non-clinical stewardship tasks (education, documentation, communication) can benefit now, whereas clinical recommendations require expert oversight. Mapping these boundaries allows AMS teams, particularly those understaffed or without on-site infectious diseases expertise, to decide where LLM support adds value rather than risk.
More Related Videos
Related Concept Videos
Clinical Significance of Antibiotic Resistance
Microbiota Modulation by Antibiotics
Antimicrobial Effectiveness
Surface Membrane Barriers
The outer layer of the skin, the epidermis, is a robust barrier comprising layers of closely packed keratinized cells. This dense arrangement prevents microbes from penetrating the body. The periodic shedding of epidermal cells...
Antibiotic Selection

