Related Experiment Video
Updated: Jul 28, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Can Generative Artificial Intelligence Reliably Score Open-Ended Question Assessments in Undergraduate Medical
Doreen M Olvet1,2, Marieke Kruidering3, Tracy B Fulton4
1Northwell, New Hyde Park, NY US.
Medical Science Educator
|July 13, 2026
Summary
Generative artificial intelligence (AI) can accurately score medical student open-ended questions (OEQs) after iterative rubric refinement. This AI-powered grading shows substantial reliability, reducing faculty workload and improving assessment efficiency.
Area of Science:
- Medical Education
- Artificial Intelligence in Assessment
- Natural Language Processing
Background:
- Open-ended questions (OEQs) are valuable for assessing medical knowledge but pose grading challenges due to time constraints.
- Generative artificial intelligence (AI) presents a potential solution for automating the scoring of OEQs.
Purpose of the Study:
- To evaluate the accuracy and reliability of generative AI (GPT-4) in scoring medical student OEQ exams.
- To determine if AI scoring can achieve comparable inter-rater reliability (IRR) to faculty grading.
Main Methods:
- Medical student OEQ responses were scored by GPT-4 using the Med2Lab platform, integrating case vignettes, questions, and rubrics.
- GPT-4 scores were compared to faculty scores using Cohen's weighted kappa (kw) for IRR.
- Iterative rubric engineering was performed based on error pattern analysis to enhance AI scoring accuracy over three iterations.
Main Results:
- Substantial IRR (kw=0.88-0.94) was achieved between faculty and GPT-4 using both analytic and holistic rubrics after three iterations.
- Moderate IRR (kw=0.54) was observed for one question using a holistic rubric.
- Score discrepancies between AI and faculty were minimal, typically only 1 point.
Conclusions:
- Generative AI, specifically GPT-4, can reliably score medical student OEQ exams.
- Iterative rubric engineering is crucial for maximizing the reliability of AI-based OEQ assessment.
- AI-powered scoring offers a promising approach to enhance efficiency in medical knowledge assessment.