Related Experiment Video
Updated: Sep 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for
Xiaoxing Gao1, Xiaoming Huang2, Rongrong Hu3
1Department of Pulmonary and Critical Care Medicine, Peking Union Medical College Hospital, 1 Shuaifuyuan, Dongcheng District, Beijing, 100730, China, 86 10 6915 5000.
Background:
Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear.
Objective:
This study aimed to assess agreement between LLM-generated and faculty ratings of history-taking and communication performance and to examine the influence of rater and case heterogeneity.
Methods:
In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, and cough). Ten blinded faculty raters scored performance (0-100 points total; 0-50 points per domain). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC [2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC).
Results:
Median total scores were similar for AI and faculty (median 93.0, IQR 89.0-95.0 vs median 94.0, IQR 91.0-95.0). Rater variability accounted for 37% of residual variance in faculty total scores (VPC=0.37). AI total scores were positively associated with faculty total scores (β=0.37, 95% CI 0.26-0.48; P<.001; Spearman ρ=0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1]=0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a small, not statistically significant mean bias (1.26 points, 95% CI -0.48 to 3.01; P=.16) and 95% limits of agreement from -4.95 to 7.48 (width=12.43 points), with proportional bias (β_proportional bias=-0.55; P<.001). Agreement was stronger for information gathering (β=0.46; ρ=0.49; ICC=0.54; VPC=0.23) than for communication (β=0.27; ρ=0.28; ICC=0.29; VPC=0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC=0.38).
Conclusions:
LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, this method is suitable for formative use and enhanced sampling in programmatic assessment but not for independent, high-stakes summative decisions.