Related Experiment Videos
What Platform Scores Miss: Multidimensional Evaluation of AI Teaching Agents in Medical Education
Hui Zhang1, Lihui Qu2, Jianmin Zheng1
1Office of Academic Affairs, Guangzhou Medical University, Guangzhou, China.
Background:
Large language model (LLM)-based AI teaching agents are increasingly used in medical education, yet their pedagogical quality is typically judged by platform-generated scores whose scoring criteria are undisclosed and may not reflect the teaching quality of the agent.
Objective:
This study aimed to develop and validate a multidimensional rubric for evaluating AI teaching agents and to examine the correspondence between platform scores and rubric-based teaching quality.
Methods:
Eight AI teaching agents covering an endocrinology curriculum were deployed across 4 role-play paradigms (patient, student, expert, and family). Twenty-two fourth-year medical students generated 167 dialogues, which were scored both by the platform and by an independently applied 8-dimension rubric (100 points, covering knowledge accuracy, pedagogical guidance, knowledge coverage, role-play quality, difficulty calibration, medical safety, student engagement, and feedback quality). Each dialogue was scored 4 times by a primary evaluator (Claude Opus 4.8; mean within-model SD 0.36), with 2 additional LLMs as robustness checks; 40 dialogues spanning all agents were rescored by a medical-education expert for validation.
Results:
Platform and rubric rankings diverged for most agents: the agent ranked third by the platform ranked last on rubric-based quality, and the platform's fourth-ranked agent ranked first. Agents differed most on knowledge-related dimensions (knowledge coverage coefficient of variation=27.3%) and least on role-play quality (coefficient of variation=5.7%), while difficulty calibration was a shared weakness. In a case-level observation, one agent revised specifically to strengthen empathy attained high role-play quality yet the lowest knowledge coverage of all agents. AI scores agreed with expert ratings at the total-score level (intraclass correlation coefficient=0.51) and on cognitive-process dimensions, but agreement was low for the more subjective dimensions. Student gender showed no detectable effect, though this analysis was underpowered.
Conclusions:
In this exploratory study, platform-generated scores reflected a construct different from agent teaching quality and should be used as a complement rather than as the sole quality indicator. The 8-dimension rubric provides a transparent, standardized alternative that reveals differences missed by platform scores, including a lack of association between empathy and knowledge coverage that warrants attention in future agent design.