Related Experiment Video
Updated: Jan 14, 2026

Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
Published on: September 27, 2020
Automated Evaluation of Reflection and Feedback Quality in Workplace-Based Assessments by Using Natural Language
Jeng-Wen Chen1,2,3,4, Hai-Lun Tu5, Chun-Hsiang Chang1,2
1Department of Otolaryngology-Head and Neck Surgery, Cardinal Tien Hospital, Fu Jen Catholic University, New Taipei City, Taiwan.
Background:
Competency-based medical education relies heavily on high-quality narrative reflections and feedback within workplace-based assessments. However, evaluating these narratives at scale remains a significant challenge.
Objective:
This study aims to develop and apply natural language processing (NLP) models to evaluate the quality of resident reflections and faculty feedback documented in Entrustable Professional Activities (EPAs) on Taiwan's nationwide Emyway platform for otolaryngology residency training.
Methods:
This 4-year cross-sectional study analyzes 300 randomly sampled EPA assessments from 2021 to 2025, covering a pilot year and 3 full implementation years. Two medical education experts independently rated the narratives based on relevance, specificity, and the presence of reflective or improvement-focused language. Narratives were categorized into 4 quality levels-effective, moderate, ineffective, or irrelevant-and then dichotomized into high quality and low quality. We compared the performance of logistic regression, support vector machine, and bidirectional encoder representations from transformers (BERT) models in classifying narrative quality. The best performing model was then applied to track quality trends over time.
Results:
The BERT model, a multilingual pretrained language model, outperformed other approaches, achieving 85% and 92% accuracy in binary classification for resident reflections and faculty feedback, respectively. The accuracy for the 4-level classification was 67% for both. Longitudinal analysis revealed significant increases in high-quality reflections (from 70.3% to 99.5%) and feedback (from 50.6% to 88.9%) over the study period.
Conclusions:
BERT-based NLP demonstrated moderate-to-high accuracy in evaluating the narrative quality in EPA assessments, especially in the binary classification. While not a replacement for expert review, NLP models offer a valuable tool for monitoring narrative trends and enhancing formative feedback in competency-based medical education.
Related Concept Videos
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Nursing Process for Patient and Caregiver Teaching III: Evaluation and Documentation
Nurses can use several methods to evaluate patient outcomes. For example, oral questions can assess cognitive learning,...
Nursing Evaluation
Self-Evaluation Maintenance Model
Self-Evaluation: Self-Enhancement and Self-Verification
Role of Communication in the Nursing Process III: Evaluation and Documentation