Related Experiment Video
Updated: Aug 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy, Reliability, and Bloom's Taxonomy Performance of Seven Large Language Models on Microbiology Questions
Volodymyr Dvornyk1, Olena Bolgova2, Volodymyr Mavrych2
1Department of Life Sciences, College of Science and General Studies, Alfaisal University, Riyadh, 11533, Saudi Arabia.
Large language models (LLMs) show strong microbiology knowledge but vary in reliability. Accuracy and consistency are key for educational use of these AI tools.
Area of Science:
- Artificial Intelligence in Medical Education
- Microbiology Knowledge Assessment
- Large Language Model Performance Evaluation
Background:
- Large language models (LLMs) are increasingly adopted in medical education.
- Systematic evaluation of LLM performance in microbiology is lacking.
- Microbiology's complex knowledge base necessitates rigorous assessment of AI tools.
Purpose of the Study:
- To benchmark seven leading LLMs on microbiology multiple-choice questions (MCQs).
- To assess LLM accuracy, test-retest reliability, and topic-specific performance.
- To analyze the impact of cognitive complexity on LLM performance.
Main Methods:
- Seven LLMs were tested on 200 microbiology MCQs across 20 topics and five Bloom's taxonomy levels.
- Testing was conducted over three independent sessions to evaluate reliability.
- Statistical analyses included ANOVA, ICCs, and Pearson correlations.
Main Results:
- Collective mean accuracy was 86.18%, with most LLMs exceeding 80%.
- Claude, Grok, and GPT demonstrated high accuracy and reliability; Gemini underperformed.
- Test-retest reliability varied significantly, with Claude showing excellent ICC (0.966) and Gemini poor ICC (0.290).
- Microbial Cell was the easiest topic (100% accuracy), Viral Genomics the most challenging (61.9%).
- Performance declined uniformly at Bloom's Level 4 (Analyze) for all LLMs.
Conclusions:
- Contemporary LLMs possess substantial microbiology knowledge but exhibit significant reliability differences.
- Response consistency is crucial for deploying LLMs in educational settings.
- Findings are specific to MCQ performance and may not extend to clinical reasoning tasks.
Related Concept Videos
Modern Molecular Taxonomy
Applications of Molecular Taxonomy
Methods to Assess Microbial Populations
Improving Translational Accuracy
Methods of Classification and Identification
Microbial Classification System
