Related Experiment Video
Updated: May 24, 2026

13:44
Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
Feasibility of LLM-Assisted Communication Feedback in Simulation-Based Medical Education: A Pilot Study
Mohamed Alhaskir1, Hannah Haven2,3, Michael Langer3,4
1Institute of Med. Informatics, Medical Faculty of RWTH Aachen University, Germany.
Studies in Health Technology and Informatics
|May 23, 2026
Summary
A new large language model pipeline provides automated feedback for pediatric medical simulations, generating reports comparable to human evaluations. This AI tool shows promise for consistent assessment in medical education.
Area of Science:
- Medical Education
- Artificial Intelligence
- Natural Language Processing
Background:
- Simulation-based medical education is crucial for training healthcare professionals.
- Automated feedback systems can enhance the efficiency and consistency of medical training.
- Assessing communication skills in pediatrics is vital for effective patient care.
Purpose of the Study:
- To develop and evaluate a local large language model (LLM) pipeline for automated feedback in pediatric simulation-based medical education.
- To assess the reliability and consistency of LLM-generated reports compared to human evaluations.
- To explore the potential of AI in standardizing communication skill assessment.
Main Methods:
- A local large language model pipeline was developed to process pediatric communication simulations.
- The system generated structured reports using the Liverpool Undergraduate Communication Assessment Scale (LUCAS).
- LLM-generated feedback was compared against human expert ratings for four simulation scenarios.
Main Results:
- The LLM pipeline produced structured reports with total scores comparable to human ratings (LLM: 13.9 ± 1.2 vs. Human: 12.8 ± 2.9).
- Scenario-level analysis showed reliable performance, though some variations were noted.
- The automated system exhibited lower inter-rater variability (SD = 0.7-1.7) than human examiners (SD = 1.3-2.9), indicating enhanced internal consistency.
Conclusions:
- The developed LLM pipeline offers a reliable method for automated feedback in pediatric medical simulations.
- The AI system demonstrates potential for consistent and objective assessment of communication skills.
- Findings support further prospective evaluation of this AI-driven educational tool.
