Related Experiment Video
Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of large language models for antimicrobial classification: implications for antimicrobial stewardship
Tan Vo1,2, Kushal Dahal1,2, Michael Klepser1,2
1College of Pharmacy, Ferris State University, 220 Ferris Dr., Big Rapids, MI, 49307, USA.
Large language models (LLMs) show promise for classifying antimicrobial medications, with top performers like Gemini and Claude achieving over 99% accuracy. Targeted feedback significantly improved LLM performance, aiding antimicrobial stewardship efforts.
Area of Science:
- Artificial Intelligence in Healthcare
- Pharmacology and Drug Classification
- Antimicrobial Stewardship
Background:
- Accurate identification of antimicrobial medications is crucial for effective antimicrobial stewardship.
- Manual classification of large medication datasets is time-consuming and prone to errors.
- Large language models (LLMs) offer potential for automating complex classification tasks.
Purpose of the Study:
- To assess the efficacy of LLMs in classifying medications as antimicrobial or non-antimicrobial.
- To evaluate the impact of targeted feedback on LLM classification accuracy.
- To determine the utility of LLMs in supporting antimicrobial stewardship programs.
Main Methods:
- A dataset of 7,239 unique medication entries was classified by four LLMs (ChatGPT-3.5, Copilot GPT-4o, Claude Sonnet 4, Gemini 2.5 Flash).
- Models underwent a two-phase evaluation: initial unguided classification followed by feedback-informed reclassification on misclassified cases.
- Performance metrics included accuracy, macro F1-score, processing time, and error reduction rates (ERRs).
Main Results:
- Post-feedback, Gemini achieved 99.6% accuracy and Claude Sonnet 4 achieved 99.4% accuracy.
- Gemini demonstrated the highest macro-F1 score (98.9%) and ERR (69.2%).
- Processing times varied significantly, with Copilot and ChatGPT-3.5 being the fastest.
Conclusions:
- High-performing LLMs can achieve accuracy levels suitable for automating initial antimicrobial classification within stewardship workflows.
- Variability in LLM performance necessitates careful model selection and ongoing human oversight for clinical applications.
- LLMs, particularly Gemini and Claude, show significant potential to enhance efficiency and accuracy in antimicrobial stewardship.
More Related Videos
11:17Multiplex Therapeutic Drug Monitoring by Isotope-dilution HPLC-MS/MS of Antibiotics in Critical Illnesses
Published on: August 30, 2018
08:58Isolation and Identification of Waterborne Antibiotic-Resistant Bacteria and Molecular Characterization of their Antibiotic Resistance Genes
Published on: March 3, 2023
Related Concept Videos
Antimicrobial Effectiveness
Development of Antibiotic Resistance
Modern Molecular Taxonomy
Microorganisms in Medicine and Therapeutics
Biological Methods for Microbial Control
Defense Against Bacterial Pathogens
Phagocytes
Phagocytes are the frontline soldiers of the immune system. They include neutrophils and macrophages. Neutrophils are the most abundant type of white blood cell and are quickly mobilized to the site of infection. Macrophages are larger cells that patrol...