Related Experiment Video
Updated: Aug 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy, Reliability, and Bloom's Taxonomy Performance of Seven Large Language Models on Microbiology Questions
Volodymyr Dvornyk1, Olena Bolgova2, Volodymyr Mavrych2
1Department of Life Sciences, College of Science and General Studies, Alfaisal University, Riyadh, 11533, Saudi Arabia.
Background:
Large language models (LLMs) are increasingly used as learning resources in medical education, yet their performance and reliability in microbiology, a discipline with a broad, heterogeneous knowledge base, have not been systematically evaluated across multiple platforms.
Objective:
To benchmark seven publicly available LLMs on microbiology multiple-choice questions (MCQs), assessing overall accuracy, test-retest reliability, topic-specific performance, and the relationship between cognitive complexity and model performance.
Methods:
Seven LLMs (Claude 4.6 Sonnet, Gemini 3.0, ChatGPT-5.2, Grok 4, Copilot, DeepSeek V3, and Kimi K2) completed 200 MCQs distributed across 20 microbiology topics and five Bloom's taxonomy levels in three independent sessions separated by 24-hour intervals. A total of 4200 responses were analyzed. Statistical analysis included one-way ANOVA with Tukey's HSD post hoc tests, repeated-measures ANOVA, intraclass correlation coefficients (ICCs), and Pearson correlations.
Results:
The collective mean accuracy was 86.18%. Six of seven systems exceeded the 80% high-competency threshold; Claude (89.83%), Grok (89.50%), and GPT (88.83%) led the group. Gemini (72.33%) was the only underperforming system. Test-retest reliability varied dramatically: Claude achieved excellent ICC (0.966), while Gemini exhibited poor reliability (ICC = 0.290), with session-to-session fluctuations of up to 100 percentage points on individual topics. Microbial Cell (100%) was the easiest topic; Viral Genomics (61.9%) was the most challenging across all systems. A uniform decline at Bloom's Level 4 (Analyze) was observed across all LLMs, with no model exceeding 78%.
Conclusion:
Contemporary LLMs demonstrate substantial knowledge of microbiology but differ markedly in reliability. Response consistency, alongside accuracy, should be a primary criterion for educational deployment. These findings are specific to microbiology MCQ performance and may not generalize to open-ended clinical reasoning.
Related Concept Videos
Modern Molecular Taxonomy
Applications of Molecular Taxonomy
Methods to Assess Microbial Populations
Improving Translational Accuracy
Methods of Classification and Identification
Microbial Classification System
