Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models for rubric-based essay grading in an undergraduate biology course
Menaka Naidu1, Nikolas S Montaquila2, Jessica P Roa2
1Brown University School of Public Health, Providence, Rhode Island, USA.
None:
This study examines how three large language models (LLMs), ChatGPT, Claude, and Gemini, assign grades to undergraduate-level essays in a biology course using a standardized rubric. Each LLM evaluated a data set of 200 essays under two prompting conditions: zero-shot (uncalibrated) and few-shot (calibrated using a small set of exemplar essays). LLM-assigned scores were directly compared with instructor-assigned scores, showing only moderate alignment with instructor grading, with variability observed across models and prompting strategies. Differences in grading behavior were also evident with different items on the rubric, with higher alignment for structural writing components and lower alignment for content- and reasoning-based criteria. Additionally, LLMs showed greater agreement with instructor scores than with one another, indicating substantial inter-model variability under identical grading conditions. These findings suggest that LLM grading outputs vary meaningfully across models, prompting strategies, and rubric components. In this context, LLMs may be best understood as tools that can support specific aspects of structured grading rather than as interchangeable evaluators.
Related Concept Videos
Improving Translational Accuracy
Reliability and Validity