Related Experiment Video
Updated: Jan 9, 2026

08:33
A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
7.4K
The Validity of Generative Artificial Intelligence in Evaluating Medical Students in Objective Structured Clinical
Masashi Yokose1, Takanobu Hirosawa1, Tetsu Sakamoto1
1Department of Diagnostic and Generalist Medicine, Dokkyo Medical University, Tochigi, Japan.
JMIR Formative Research
|December 4, 2025
Summary
Generative AI like ChatGPT-4 shows potential to assist in Objective Structured Clinical Examinations (OSCE), but agreement with physician evaluations was poor in most domains. Further research is needed to confirm its validity as a complementary OSCE assessor.
Area of Science:
- Medical Education Technology
- Artificial Intelligence in Healthcare
- Clinical Skills Assessment
Background:
- Objective Structured Clinical Examination (OSCE) is a standard medical education assessment tool.
- OSCE implementation is resource-intensive, posing challenges for medical institutions.
- Generative AI, such as ChatGPT-4, is explored as a potential solution to reduce physician assessment burden.
Purpose of the Study:
- To evaluate the validity of generative AI (ChatGPT-4) as a complementary assessor for OSCE.
- To compare evaluation scores assigned by generative AI and human physicians.
- To assess the agreement between AI and physician evaluations in OSCE.
Main Methods:
- An experimental study involved 11 fifth-year medical students performing a mock patient interview.
- Physicians and ChatGPT-4 evaluated student performance using a 6-domain rubric and Likert scale.
- Statistical analysis included the Wilcoxon signed-rank test and intraclass correlation coefficients (ICCs).
Main Results:
- ChatGPT-4 assigned higher scores than physicians in physical examination, patient notes, clinical reasoning, and management domains.
- No significant score differences were observed in patient care/communication and history taking.
- Intraclass correlation coefficients (ICCs) indicated poor agreement between ChatGPT-4 and physician evaluations across most domains.
Conclusions:
- Generative AI (ChatGPT-4) shows potential to support OSCE assessment in specific domains.
- The poor agreement necessitates further research to establish AI's reproducibility and validity in OSCE.
- ChatGPT-4 may serve as a complementary tool, but not a replacement for physician assessment in OSCE.
