Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for
Xiaoxing Gao1, Xiaoming Huang2, Rongrong Hu3
1Department of Pulmonary and Critical Care Medicine, Peking Union Medical College Hospital, Beijing, CN.
Background:
Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear.
Objective:
To assess agreement between LLM-generated and faculty ratings of history-taking and communication performance, and to examine the influence of rater and case heterogeneity.
Methods:
In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, cough). Ten blinded faculty raters scored performance (0-100 total; 0-50 domains). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC[2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC).
Results:
Median total scores were similar for AI and faculty (93.0 [IQR 6.0] vs 94.0 [IQR 4.0]). Rater variability accounted for 37% of residual variance in faculty total scores (VPC = 0.37). AI total scores were positively associated with faculty total scores (β = 0.37, 95% CI 0.26-0.48, P<.001; Spearman ρ = 0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1] = 0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a non-significant mean bias (1.26 points) and 95% limits of agreement from -4.95 to 7.48 (width = 12.43 points), with proportional bias (β_mean = -0.55, P<.001). Agreement was stronger for information gathering (β = 0.46, ρ = 0.49, ICC = 0.54, VPC = 0.23) than for communication (β = 0.27, ρ = 0.28, ICC = 0.29, VPC = 0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC = 0.38).
Conclusions:
LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, it is suitable for formative use and enhanced sampling in programmatic assessment, but not for independent high-stakes summative decisions.
