Related Experiment Video
Updated: Jul 11, 2026

05:33
Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning
Published on: January 29, 2020
Agreement Between Reasoning-Oriented Generative AI Models and Clinical Educators in Evaluating Japanese Objective
Takanobu Hirosawa1, Masashi Yokose1, Tetsu Sakamoto1
1Department of Diagnostic and Generalist Medicine, Dokkyo Medical University, 880 Kitakobayashi, Mibu-cho, Shimotsuga, Tochigi, 321-0293, Japan, 81 282-87-2498.
JMIR Formative Research
|July 2, 2026
Summary
Generative artificial intelligence (GenAI) models showed poor agreement and lower scores compared to clinical educators for evaluating Japanese medical interviews. These AI tools are not yet suitable as standalone evaluators for Objective Structured Clinical Examination transcripts.
Area of Science:
- Medical education technology
- Artificial intelligence in healthcare
- Clinical skills assessment
Background:
- Medical interview training evaluation faces implementation and assessment challenges.
- Generative artificial intelligence (GenAI) presents a potential solution, but its effectiveness in Japanese language evaluations is uncertain.
Purpose of the Study:
- To assess scoring patterns and evaluate agreement between reasoning-oriented GenAI models and clinical educator consensus ratings.
- To determine the utility of GenAI for evaluating Japanese medical interview training.
Main Methods:
- Two blinded clinical educators reached consensus ratings on 40 Japanese medical interview transcripts from residents.
- Two GenAI models (GPT-5.2 Thinking and Gemini 3.0 Pro) independently evaluated the same transcripts using a standardized 6-domain rubric.
- Evaluations were compared using the Wilcoxon signed-rank test, and interrater reliability was assessed using intraclass correlation coefficients (ICCs).
Main Results:
- Clinical educator consensus ratings showed the highest mean scores (5.18).
- GenAI models scored significantly lower: GPT-5.2 Thinking (3.68) and Gemini 3.0 Pro (4.09).
- Agreement between GenAI models and clinical educators was poor (ICCs: 0.04 for GPT-5.2 Thinking, 0.22 for Gemini 3.0 Pro).
Conclusions:
- Preliminary findings indicate GenAI models exhibit lower scores and poor agreement with expert consensus for Japanese medical interview transcripts.
- Current GenAI models, under tested conditions, are not recommended as standalone evaluators for Objective Structured Clinical Examination assessments in Japanese.
- Further research is needed to explore GenAI's potential for providing formative feedback in medical education.