Related Experiment Video
Updated: Jun 10, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparing Scoring Consistency of Large Language Models with Faculty for Formative Assessments in Medical Education
Radhika Sreedhar1, Linda Chang2, Ananya Gangopadhyaya2
1University of Illinois College of Medicine, Chicago, IL, USA. sreedhar@uic.edu.
Large language models (LLMs) show promise in assessing medical student critical appraisal assignments, offering consistent feedback and reducing faculty time. This study found LLMs comparable to faculty grading, highlighting their potential in medical education.
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
- Educational Technology
Background:
- Medical education requires individualized feedback on self-directed learning skills for pre-clinical students.
- Critical appraisal assignments are used, but faculty feedback is time-intensive.
- Large language models (LLMs) offer potential for automated scoring and feedback generation.
Purpose of the Study:
- To evaluate the consistency and feasibility of using LLMs for assessing and providing feedback on undergraduate medical student formative assessments.
- To explore the psychometric characteristics of LLM-based assessment.
Main Methods:
- Cross-sectional study of 111 pre-clinical medical student critical appraisal assignments.
- Comparison of ChatGPT 3.5 scoring against existing faculty grades using a developed prompt.
- Analysis of scoring differences, inter-rater reliability (IRR), internal-consistency reliability, area under precision-recall curve (AUCPR), and cost-effectiveness.
Main Results:
- LLM scoring of individual items was comparable to faculty grading.
- Overall agreement between LLM and faculty was 67% (P < 0.001).
- LLM use reduced faculty time fivefold, potentially saving 150 faculty hours.
Conclusions:
- LLMs demonstrate potential as a tool to assist faculty in assessing formative assignments and providing feedback in medical education.
- The study highlights the psychometric validity and feasibility of using LLMs for educational assessment.
More Related Videos
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Reliability and Validity
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...