Related Experiment Video
Updated: Jun 27, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Validation of the Use of a Large Language Model for Detecting Sentiment in Student Course Evaluation
Kate Rowland1, Ling Wang2, Kirstie Bash3
1Rush University Medical Center, Chicago, IL.
Family Medicine
|May 8, 2026
Summary
A bidirectional encoder representations from transformers (BERT) model can effectively identify sentiment in medical student evaluations. This artificial intelligence tool demonstrated comparable reliability to human coders, supporting its use in medical education.
Area of Science:
- Medical Education
- Natural Language Processing
- Artificial Intelligence
Background:
- Large language models and NLP are increasingly used in medical education.
- Risks of bias and errors necessitate validation of AI tools before research or educational use.
- This study focused on validating a specific NLP method for analyzing student feedback.
Purpose of the Study:
- To validate the application of a bidirectional encoder representations from transformers (BERT) model.
- To identify the presence and patterns of sentiment in end-of-course evaluations from medical school clerkships.
- To assess the reliability of AI-driven sentiment analysis compared to human coders.
Main Methods:
- Utilized the Patino framework for AI validation in health professions education.
- Human coders analyzed de-identified course evaluation comments, calculating human-human interrater reliability.
- Trained a BERT model using human-identified keywords and calculated human-AI interrater reliability.
Main Results:
- 364 comments were analyzed, with sentiment distribution (positive, negative, neutral, mixed) varying across institutions.
- Sentiment patterns and reliability metrics differed by school.
- Human-AI interrater reliability was comparable to human-human interrater reliability.
Conclusions:
- Conceptual frameworks exist for validating AI tools in health professions education.
- A trained BERT model can accurately detect sentiment in medical student evaluations.
- The BERT model's reliability is similar to human coders, suggesting its utility in educational assessment.