Related Experiment Video
Updated: Jun 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
LLM-based automatic short answer grading in undergraduate medical education
1Department of Life Sciences and Medicine, University of Luxembourg, 6, avenue de la Fonte, L-4364, Esch-sur-Alzette, Luxembourg. christian.grevisse@uni.lu.
Large Language Models (LLMs) show promise for automatic short answer grading (ASAG) in medical education. Gemini 1.0 Pro
Area of Science:
- Medical Education
- Artificial Intelligence
- Natural Language Processing
Background:
- Traditional multiple-choice questions in medical education favor recognition over recall.
- Grading open-ended questions is time-consuming for educators.
- Automatic short answer grading (ASAG) offers a solution, with recent advancements in Large Language Models (LLMs) driving progress.
Purpose of the Study:
- To evaluate the efficacy of LLMs, specifically GPT-4 and Gemini 1.0 Pro, for automatic short answer grading in undergraduate medical education.
- To compare LLM grading performance against human evaluators.
Main Methods:
- 2288 student answers from 12 undergraduate medical courses across 3 languages were graded.
- GPT-4 and Gemini 1.0 Pro were utilized for the grading process.
- LLM grades were compared with human evaluator grades.
Main Results:
- Gemini 1.0 Pro's grades closely matched human evaluators, while GPT-4 assigned lower grades but had fewer false positives.
- Both LLMs demonstrated moderate agreement with human grades and high precision for fully correct answers with GPT-4.
- LLM grading consistency was observed for high-quality answer keys, with weak correlations to answer length or language.
Conclusions:
- LLM-based ASAG requires human oversight in medical education but can save educators time on grading edge cases.
- The inherent knowledge of LLMs appears sufficient for Bachelor-level medical education, negating the need for fine-tuning.
- LLMs can assist educators by automating grading, allowing more focus on complex student responses.
More Related Videos
07:32Use of Galvanic Skin Responses, Salivary Biomarkers, and Self-reports to Assess Undergraduate Student Performance During a Laboratory Exam Activity
Published on: February 10, 2016
04:12Mixed Reality for Education MRE Implementation and Results in Online Classes for Engineering
Published on: June 23, 2023
Related Concept Videos
Surveys
Reliability and Validity