Related Experiment Video
Updated: Sep 12, 2025

Nest Building Behavior as an Early Indicator of Behavioral Deficits in Mice
Published on: October 19, 2019
Artificial intelligence assisted automated short answer question scoring tool shows high correlation with human
H M T W Seneviratne1, S S Manathunga2
1Department of Pharmacology, Faculty of Medicine, University of Peradeniya, Peradeniya, Sri Lanka. thilanka.medi@gmail.com.
An AI-powered automated short answer question scoring tool (ASST) shows high accuracy in grading medical student responses. This artificial intelligence (AI) system provides reliable, granular feedback, correlating strongly with human examiner marks.
Area of Science:
- Medical Education
- Artificial Intelligence in Education
- Natural Language Processing
Background:
- Optimizing short answer question (SAQ) assessment and feedback for medical undergraduates is challenging due to increasing student numbers and staff shortages.
- Developing scalable and effective assessment tools is crucial for medical education.
- Automated grading systems offer a potential solution to these challenges.
Purpose of the Study:
- To develop and evaluate an automated SAQ scoring tool (ASST) utilizing artificial intelligence (AI).
- To assess the feasibility of using large language models (LLMs) for grading SAQs based on provided rubrics.
- To provide personalized, granular feedback to students on their written answers.
Main Methods:
- Investigated the use of GPT-4, a large language model (LLM), for automated SAQ scoring in a Systematic Pharmacology course.
- LLM analyzed student answers against provided rubrics, extracting key information, scoring, and generating feedback.
- Evaluated LLM performance by averaging five sampled runs and compared scores against human examiner assessments for correlation.
Main Results:
- The automated SAQ scoring tool (ASST) demonstrated a high correlation with human examiner markings, with correlation coefficients of 0.93 and 0.96 across 30 student answers.
- Excellent inter-rater reliability was observed, with an intra-class correlation coefficient of 0.94 between the LLM and human graders.
- The AI-assisted tool provided scores that closely matched expert human judgment.
Conclusions:
- AI-assisted automated SAQ scoring tools show significant promise for transparent and flexible grading in medical education.
- The system's high correlation with human graders suggests its potential to reduce instructor workload.
- This approach offers students more granular and actionable feedback, enhancing the learning process.
More Related Videos
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
05:21Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
Related Concept Videos
Non-equilibrium in the Cell
Microsoft Excel: Pearson's Correlation