Related Experiment Video
Updated: May 27, 2026

05:22
Systematic Endobronchial Ultrasound - The Six Landmarks Approach
Published on: August 11, 2023
Large language model-generated versus teacher-written objective structured clinical examination stations for medical
Piotr Szychowiak1, Jonathan Wong-So1, Hélène Messet1
1Médecine Intensive Réanimation, Centre Hospitalier Universitaire d'Orléans, Orleans, France.
Summary
Large language models (LLMs) show potential for creating objective structured clinical examination (OSCE) stations, but GPT-4o requires significant teacher review for reliable assessment grids and overall quality.
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
Background:
- Developing Objective Structured Clinical Examination (OSCE) stations is a resource-intensive task for medical educators.
- There is a growing interest in leveraging artificial intelligence, specifically large language models (LLMs), to streamline educational content creation.
Purpose of the Study:
- To evaluate the efficacy of the LLM GPT-4o in generating ready-to-use OSCE stations for medical education.
- To compare the quality of LLM-generated OSCE stations against those developed by experienced medical teachers.
Main Methods:
- Five OSCE stations were generated using GPT-4o, with reference knowledge provided.
- Seven expert assessors evaluated both LLM-generated and teacher-written stations using a 5-point Likert scale.
- Stations were assessed for overall quality and readiness for student use.
Main Results:
- All five teacher-written stations were deemed high quality.
- Only one of the five GPT-4o-generated stations met the quality threshold.
- GPT-4o generated adequate clinical scenarios but struggled with creating reliable assessment grids.
Conclusions:
- LLMs like GPT-4o can assist in generating clinical scenarios for OSCEs, but require substantial expert oversight.
- Current LLM capabilities are insufficient for producing fully reliable and ready-to-use OSCE stations independently.
- Medical teachers' careful review and refinement remain critical for ensuring the quality and validity of AI-generated OSCEs.