Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for

Xiaoxing Gao1, Xiaoming Huang2, Rongrong Hu3

  • 1Department of Pulmonary and Critical Care Medicine, Peking Union Medical College Hospital, Beijing, CN.

Summary

Large language model (LLM)-powered virtual standardized patients (VSPs) show moderate agreement with faculty in scoring clinical skills. LLM scoring is better for information gathering than communication, suitable for formative assessment.

Related Concept Videos