Related Experiment Video
Updated: Sep 26, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Automated grading of short-answer image-based assessments using a hybrid natural language processing-AI framework:
Xu Hao Isaac Tan1, Zhen Li Samantha Lee1, Thida Win1
1Department of Radiology, KK Women's and Children's Hospital, Singapore.
Introduction:
Efficient and reliable assessment is essential in medical education. Manual grading of short-answer radiology quizzes is resource-intensive and variable. Automated marking systems can alleviate these challenges. This study introduces a novel automated grading system leveraging techniques in natural language processing (NLP) and generative artificial intelligence. Our objective was to assess the system's agreement against human grading and to analyse how response length influences grading concordance.
Methods:
We obtained 840 responses from 28 radiology residents to 30 X-ray quiz questions. The system employed string-matching algorithms and rule-based logic for anatomical and laterality mismatches. A ChatGPT-created synonym dictionary was used to recognise acceptable alternative answers. Cohen's κ was used to assess inter-rater agreement across answer lengths, and logistic regression was used to examine the response length's relationship with grading disagreement.
Results:
To establish a reference standard, two independent human raters graded the responses, achieving near-perfect inter-rater reliability (κ = 0.985). Evaluated against human consensus, the system achieved 97.5% raw agreement (816/837). Perfect agreement (κ = 1.00) was observed for one-word responses (n = 429), with near-perfect agreement for two-word responses (κ = 0.927, n = 75). Agreement decreased as response length increased, with each additional word increasing the odds of disagreement by 51.9% (odds ratio 1.519; 95% confidence interval 1.227-1.879).
Conclusion:
The automated grading system, enhanced by NLP techniques and ChatGPT-assisted synonym expansion, demonstrated strong concordance with human grading, particularly for one- and two-word answers. This hybrid approach offers a scalable solution for short-answer assessment, with strong potential in radiology and other medical education contexts.