Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for
Xiaoxing Gao1, Xiaoming Huang2, Rongrong Hu3
1Department of Pulmonary and Critical Care Medicine, Peking Union Medical College Hospital, Beijing, CN.
JMIR Medical Education
|July 31, 2026
Summary
Large language model (LLM)-powered virtual standardized patients (VSPs) show moderate agreement with faculty in scoring clinical skills. LLM scoring is better for information gathering than communication, suitable for formative assessment.
Area of Science:
- Medical education technology
- Artificial intelligence in healthcare
- Clinical skills assessment
Background:
- Large language model (LLM)-powered virtual standardized patients (VSPs) offer scalable clinical skills practice.
- The validity of AI-generated scores compared to faculty ratings is not well-established.
Purpose of the Study:
- To evaluate the agreement between LLM-generated and faculty ratings of medical students' history-taking and communication skills.
- To analyze how rater and case variability influence assessment agreement.
Main Methods:
- 92 fourth-year medical students participated in voice-based VSP cases (fever, diarrhea, cough).
- Ten faculty raters and an LLM (DeepSeek-V3) scored student performance.
- Agreement was analyzed using mixed-effects models, ICC, Spearman correlations, MAE, and Bland-Altman plots.
Main Results:
- LLM and faculty median total scores were comparable (93.0 vs 94.0).
- Moderate absolute agreement was observed (ICC=0.51), with a mean absolute error of 3.11 points.
- Agreement was higher for information gathering (ICC=0.54) than for communication (ICC=0.29).
Conclusions:
- LLM-based scoring in VSPs demonstrates moderate agreement with faculty ratings.
- LLM performance is superior in evaluating information gathering compared to communication skills.
- LLM scoring is appropriate for formative assessment and enhanced sampling, not high-stakes summative decisions.
