Related Experiment Video
Updated: Jul 9, 2026

08:20
Superior Auto-Identification of Trypanosome Parasites by Using a Hybrid Deep-Learning Model
Published on: October 27, 2023
From setting to vetting: using artificial intelligence for Single Best Answer questions review
Olivia Ng1, Siew Ping Han1, Magdalene Hui Min Lee2
1Lee Kong Chian School of Medicine, Nanyang Technological University, Singapore.
Korean Journal of Medical Education
|March 9, 2026
Summary
Artificial intelligence (AI) can assist in vetting Single Best Answer (SBA) questions by acting as a first-pass reviewer, enhancing consistency and reducing workload in medical education. Human oversight remains crucial for ensuring educational relevance.
Area of Science:
- Medical Education Technology
- Artificial Intelligence in Assessment
- Educational Measurement
Background:
- Maintaining the quality of Single Best Answer (SBA) questions is a persistent challenge in medical education.
- The rise of AI-generated assessment items necessitates robust vetting processes.
- Current AI question generation is well-studied, but AI-assisted vetting remains underexplored and difficult to scale.
Purpose of the Study:
- To investigate the feasibility and reliability of using a large language model (LLM) to support the vetting of SBA questions.
- To develop and evaluate an AI-based reviewer (QA-bot) for SBA question quality assessment.
- To compare AI reviewer performance against experienced human educators.
Main Methods:
- Developed an AI reviewer (QA-bot) using custom GPT, incorporating 25 criteria based on Bloom's taxonomy (Levels 1-3).
- Utilized a shared evaluation rubric for independent assessment of 32 AI-generated SBA questions.
- Assessed inter-rater reliability between QA-bot and human reviewers using intraclass correlation coefficients (ICC).
Main Results:
- The evaluation rubric demonstrated high internal consistency (Cronbach's alpha=0.878) and strong inter-rater reliability among human reviewers (ICC=0.893).
- QA-bot showed good alignment with human raters (ICC=0.861 and 0.840).
- The AI performed well on objective criteria but was less consistent in identifying irrelevant complexity and judging question difficulty.
Conclusions:
- AI can serve as an efficient initial reviewer for SBA questions, improving consistency and decreasing workload.
- Human oversight is indispensable for ensuring the educational and clinical relevance of assessment items.
- AI-assisted vetting offers a scalable solution to enhance the quality assurance of medical education assessments.