Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing multiple-choice question quality in internal medicine: a comparative analysis of three large language
Mevlüt Okan Aydin1, Belkıs Nihan Coşkun2, İbrahim Hamal3
1Department of Medical Education, Faculty of Medicine, Bursa Uludağ University, Bursa, Türkiye.
Background:
Large language models (LLMs) are increasingly explored for their potential to support quality assurance in medical education assessment. However, limited evidence exists on the alignment between LLM evaluations and expert judgment across multiple dimensions of multiple-choice question (MCQ) quality.
Methods:
This comparative methodological study evaluated 85 MCQs from an internal medicine clerkship examination. Three LLMs (Claude Sonnet 4, Gemini 2.5 Flash, and Llama 3.3 70B Instruct Turbo) and three medical education experts independently assessed each question for cognitive level (Revised Bloom's Taxonomy), alignment with the intended learning outcome (5-point Likert scale), and presence of technical flaws based on NBME guidelines. Agreement was calculated using Fleiss' kappa for cognitive level classification and Cohen's kappa for binary technical flaw criteria, with intraclass correlation coefficients (ICC) for Likert-scale alignment ratings.
Results:
For cognitive level classification, Gemini (κ = 0.424, p < 0.001) and Claude (κ = 0.415, p < 0.001) showed moderate agreement with experts; Llama demonstrated lower agreement (κ = 0.266, p < 0.001). Alignment with learning outcomes yielded weak-to-moderate agreement for all models (ICC 0.121-0.380). For technical adequacy, Claude and Gemini achieved almost perfect agreement on detecting negatively worded stems (κ = 0.897, p < 0.001) and substantial agreement on inconsistent numerical data (κ = 0.661, p < 0.001) but showed poor agreement on more subjective flaws. Llama performed poorly across most technical criteria.
Conclusion:
Claude and Gemini demonstrate moderate to strong agreement with experts for cognitive level classification and detection of objective technical flaws, suggesting their potential as adjunctive tools in MCQ review. However, weak agreement on learning outcome alignment and variability across models indicates that LLMs cannot yet replace expert judgment. A hybrid approach combining LLM-assisted screening with human expertise may optimize item quality assurance in medical education. These findings derive from a single institution and discipline with a limited item set (n = 85) and require multi-center validation before broader generalization.