Related Experiment Videos
Development and External Validation of a Machine Learning Model for Automated Feedback Quality Assessment in Chinese
Lifeng Yao1, Yijun Chen1, Jing Shen1
1Department of Anesthesiology, The First Affiliated Hospital of Ningbo University, Ningbo, Zhejiang, People's Republic of China.
Advances in Medical Education and Practice
|May 7, 2026
Summary
This study developed a machine learning model to automate the evaluation of narrative feedback quality in anesthesiology residency programs. The model shows promise as a screening tool for competency-based medical education.
Area of Science:
- Medical Education
- Machine Learning
- Anesthesiology
Background:
- High-quality narrative feedback is crucial for competency-based medical education.
- Manual evaluation of feedback is time-consuming and subjective.
- Automating feedback assessment can improve efficiency and objectivity.
Purpose of the Study:
- To develop and validate a machine learning (ML)-based model for automated evaluation of narrative feedback quality.
- To assess the model's performance in the context of anesthesiology residency training.
- To support competency-based medical education through efficient feedback analysis.
Main Methods:
- Utilized 990 narrative feedback entries for training and validation, with 587 for external testing.
- Employed TF-IDF and manual features with Chinese word segmentation and an anesthesia-specific vocabulary.
- Addressed data imbalance using SMOTE and compared Logistic Regression (LR), Random Forests (RF), and Gradient Boosting Machine (GBM).
Main Results:
- The LR model achieved optimal internal performance (F1 score: 0.941, cross-validation accuracy: 0.925 ± 0.026).
- Externally, the LR model showed 0.840 accuracy, high recall (0.956), and moderate precision (0.636) for high-quality feedback (F1: 0.764, AUC: 0.729).
Conclusions:
- Successfully developed and externally validated an ML model for automated feedback quality assessment in Chinese anesthesiology residency.
- The model demonstrates high recall and stable internal performance, suitable for batch evaluation.
- The model can serve as a screening tool to enhance competency-based medical education.